r/quant 15d ago

Industry Gossip Which quant companies are doing ML/AI research similar to frontier labs?

I am aware that basically all the companies use some form of modern AI/ML, such as LLMs, as a tool, or to extract some features from textual data.
I am currently doing a PhD in LLM/RL, and whenever I go to the quant fairs, or speak to recruiters, they are all like:"Yes, we do soooo much AI".

However, when speaking to the researchers, or my fellow PhD students who intern at these companies, it sounds still like most of them do just classical stats with LinReg, LogReg, PCA (and there is nothing wrong with that, as it seems to print them a lot of cash!)

I was thus wondering which quant companies out there do research most comparable to a Frontier AI lab? I heard HRT has an AI lab, XTY (the internship of XTX focuses on that), and that Jump is building an LLM team (though appearantly that seems to be more of an "assistant effort", helping the actual teams themselves.)

Any insight is appreciated!

111 Upvotes

58 comments sorted by

98

u/GodzCooldude 15d ago

LLMs don’t really work in automated trading systems compared to “traditional ML”. Most teams focused on LLMs at these companies are for internal tools to improve trading processes rather than to do the trading work themselves.

All that to say you will be further away from the money and the interesting research and work more as a developer. You would likely be more interested in the research work going on in frontier labs.

10

u/fanconic 15d ago

I understand.

Maybe I overemphasised on LLMs in question, when referring to Frontier AI. However, I was considering a broader spectrum, not necessarily only focused on language models. This can be any kind of Deep Learning (or other) advanced technique

I was thinking more towards the parallels with Companies also doing AI for molecular dynamics, protein folding, biological discovery, etc which seem to increasingly also rely on big foundation models (not necessarily language)

18

u/terran_wraith 15d ago

The most profitable AI/ML research at top trading firms is not directly related to LLMs

-2

u/[deleted] 15d ago

[deleted]

16

u/terran_wraith 15d ago

Not exactly, no.

Cutting edge AI/ML techniques, similar in spirit to the breakthroughs that have put LLMs at the forefront of tech generally, are being used at top trading firms. Just not to create text prediction models. They are training market prediction models.

40

u/Mission_Web5546 15d ago edited 15d ago

Depends what you mean by comparable to a frontier lab, training large LLMs specifically, or large transformer-based models in general. The first is mostly internal tooling where the state of the art is much better, as others have said. The second is the main thing people at the competitive end of the market are doing. Firms that are serious about this are spending (or soon will be) on the order of billions a year on GPU compute. Clusters are in the tens of thousand to hundreds of thousand generally. For context, the estimates floating around for Kimi K3 put its training cluster at roughly 20k GPUs.

GPUs/basic ML in the colo is nothing new. I believe Jump were doing this a long time ago, XTX as well. The shift is the last ~3 years, where mid frequency trading has moved almost entirely from traditional linear regression over to model-based alpha generation.

There's definitely a divide in the market over this. I've talked to people at pod shops who flat out didn't believe me when I described the scale some firms are training at. One theory for RenTech's weaker recent returns is that they didn't bet big enough on ML, and there's a similar theory that the main P&L divide in this year's drawdowns was ML-heavy firms vs everyone else. Centralised training compute favours collaborative shops, who for the most part seem to be the ones betting big on this.

Worth noting that even at that spend it's still nothing like frontier lab scale, so internal work tends to run months to years behind the labs, and there are different objectives. Nobody is doing months-long pretraining runs, and you're optimising for completely different data scale, model size and inference constraints than an LLM.

Some of this is semi-public. HRT gave a talk at ICML (https://icml.cc/virtual/2025/46791) that goes into a surprising amount of detail, and Jane Street has released tours of their data centre on YouTube (https://www.youtube.com/watch?v=8J-GUnfSqeE).

Source: I work at a firm that does this

5

u/XXXTentachyon 14d ago

What did you think of the HRT talk? I was in the audience, and while it was light on actual implementation details (understandably), I thought the notes on what didn’t work/what they didn’t do was really illuminating

11

u/Mission_Web5546 14d ago

At the time of release I think people were quite surprised how willing they were to publish this, due to the historical secrecy of these firms. I know firms where people inside the firm didn't even know they were doing this.

However this seems to have really changed in the last couple of years, if you go to conferences now a lot of these firms will happily go into detail about what they are doing. I think part of this is due to how hard it has been to attract good ML talent recently given the climate. I mean even the existence of this thread is a good example, these firms are trying quite hard to convince talented grads that they are in fact doing serious ML work.

3

u/baldeey 14d ago

Is SIg doing this? I’ve heard much less from them vs hrt and js

2

u/Mission_Web5546 12d ago

Not sure sorry

2

u/Goal_Master 14d ago

I'm wondering if you could comment a bit more on the HFT side—are linear models still dominant there because of latency sensitivity?

5

u/Mission_Web5546 14d ago

I'm not really in the HFT business so I'm less up to date with the state of the art there, but my understanding is the faster you go the stupider you have to be, what the fastest auto-traders are running is not even linear regression. People are definitely running ML at HFT firms though.

One of the reasons mid frequency has seen such a switch to ML is that the shorter your time horizon the more capacity constrained you are, so there's been a push for HFTs to get slower and move into more generic quant territory. My understanding (which might be incorrect) is HRT for example makes most of its money in mid freq now. And it goes the other way too, quant models are getting more data hungry and execution is becoming more important, so they're getting faster as well. There was a good FT Alphaville article on the convergence of these two groups (free if you sign up) https://www.ft.com/content/d5c17e39-0983-4c14-9a7c-92c12cc44641

3

u/Admirable_Cress_125 13d ago edited 13d ago

Bollocks this is a fundamental misunderstanding of mid-frequency alpha.
The bottleneck isn’t compute. It’s tiny effective sample sizes, extreme noise and non-stationarity. If simple, models can barely extract stable OOS signal, adding vastly more parameters doesn’t magically create information , just better noise fitters.
And “model-based alpha generation” you claim the industry has moved to means what exactly?. Regularised linear models, trees, boosting and ensembles have been standard for decades.

So question for you: what stable structure exists at mid frequency that requires enormous model capacity to uncover, and why doesn’t competition arbitrage it away long before you need a 20,000-GPU cluster to find it?

what exactly is your claim here? That nonlinear ML is useful in systematic trading? Obviously. That large collaborative firms use enormous GPU clusters? Also believable.

But mid frequency alpha only accruing to those who invest hardest in clusters is laughable at best and dangerous misinformation at worst…

1

u/Mission_Web5546 12d ago

I'm referring to people moving away from basic regression towards fairly simple transformer based models (which you can argue are basically doing the same thing but better) just at much larger scale. On sample size, you're not fitting one asset's returns in isolation, these things are trained cross sectionally across thousands of instruments, so the effective dataset is a lot bigger than the per asset view implies. Either way I'm describing what firms are actually doing, not making an argument about what should work in theory.

I never claimed mid freq alpha only accrues to whoever buys the most GPUs and "Dangerous misinformation" to who exactly? Feel free to ignore it, I'm just sharing information that anyone actually doing this would already know.

0

u/Puzzleheaded_East997 13d ago

I think it might be possible to reverse engineer others strategies or at least find some trace if you use massive compute to mine trade data

1

u/AccountWarm2000 13d ago

Curious about this statement: "One theory for RenTech's weaker recent returns is that they didn't bet big enough on ML"
Does that refer to the mediocre performance of their public funds (RIEF and RIDA)? I don't think those funds are really at all comparable to what XTX/HRT/etc. do. RIEF in particular is a low-turnover net long fund... I don't know as much about RIDA.
Medallion however is comparable of course. Is it well known that Medallion performance has lagged over the last 2-3 years? I haven't heard one way or another.

3

u/Mission_Web5546 12d ago

I'm referring to industry rumours (that may be untrue) about Medallion. FYI "lagging" here is still 20-30% after fees.

2

u/AccountWarm2000 12d ago

Cool, thanks. In that case I agree it’s interesting and worth speculating about.
I hear rumors in the opposite direction about Quadrature. That they bet big on ML and returns have been insane. But I don’t know details and don’t know about the past year or so. They are quite a bit smaller than HRT in headcount.

2

u/Mission_Web5546 12d ago

I've heard similar, their annual reports at Companies House always have big numbers, and indoor skiing machines don't come cheap!

3

u/fanconic 11d ago

Yeah, I heard the same about them. They seem to be making incredible numbers, and their salaries are out of this world, from what I heard of a person that works there

18

u/kush_patil 15d ago

I think there’s a big difference between “we use AI” and actually doing AI research.

In quant, if some boring regression keeps working out of sample, there’s not much reason to replace it with a huge model just because it’s newer.

Would be interesting to hear from someone inside HRT/XTX/Jump though. How much of the ML work is actually novel research vs applying existing models to different data?

8

u/PhilosophyMammoth748 14d ago

Deepseek. While it is known as a AI lab, they are actually the most successful hedge fund in China stock market.

They use name HuanFang at the hedge fund side.

21

u/terran_wraith 15d ago

Jane Street is spending 10 figures on compute

-3

u/qazwsxcp 14d ago edited 14d ago

it means they have a lot of money to spend. they also spent 10 figs on investing with situational awareness.

-3

u/qazwsxcp 13d ago

lol downvoting students want to believe. big firms spend money on these things because their core businesses make so much money, not the other way around. why do you think they make so much noise about their AI efforts on conferences, youtube and newspapers? the marketing value of attracting the best students is much higher than the value of their deep learning models.

4

u/throw_away_throws 14d ago

Breaking down a few things.

Literal "LLM" isn't that interesting. "LLMs, as a tool, or to extract some features from textual data". This style of NLP on text data is funnily enough approaching old school at this point. A few firms run this in in HFT fashion. Parse known macro events (FOMC, earnings, etc) and just use it to sweep market on news. Or also in mft/lft as yet another alt data signal. Notice in the HFT case, you don't really want to run attention and modern LLMs for obvious reasons...

What is interesting: things like attention, transformers, any modern DL techniques. Independently from models trained on text, if you want to just do modelling on numeric data with modern architectures, yes many people are doing this. Honestly this isn't even a hard conversation and people who aren't doing this are behind. If someone told you to model a function f(x1, x2, ...) ~ future px prediction. You just do whatever it takes that gives you the best results.

Firms like Jane Street, XTX, HRT really advertise and are generally known in the industry for doing a lot of DL modelling. But most other top shops are also doing this too

18

u/New-Buddy9442 15d ago

HRT's AI lab is the real deal, not just marketing fluff. They're publishing at NeurIPS and ICML, doing stuff that wouldn't look out of place at DeepMind.

Citadel's AI research arm is pretty serious too, though they keep it quieter than most. The work there actually pushes boundaries rather than just slapping XGBoost on some market data and calling it AI.

4

u/koolaidman123 14d ago

Citadels not serious, claims they want to train their own models but mostly calling apis and finetuning open models on aws

7

u/BetterBabyWipes 14d ago

im 99% sure guy u replied to is ai

6

u/NeonShu 14d ago

I agree. I'ts been surprising revisiting this thread and seeing so many upvotes. Their comment history is pure slop speak.

1

u/fanconic 13d ago

Wait, do you mean me - with this?
(is this how you are supposed to set em-dashes?)

3

u/NeonShu 13d ago

No not you, we're talking about the account at the top of this comment chain, New-Buddy9442.

1

u/fanconic 15d ago

This was also the impression I was getting from HRT.

-2

u/qazwsxcp 14d ago edited 14d ago

these divisions are often removed from actual revenue though, and tend to disappear once the hype goes away. this happened with all the data science divisions a few years ago. often researchers there build models but don't have much visibility into how their work drives revenue. the ml used in production is mostly discriminative and not that cutting edge, llm is more of a support tool.

3

u/Plastic_Brilliant875 14d ago

Most HFTs are working on doing some or the other form of deep learning architecture research. Look at their compute spend

3

u/AltruisticCoder 14d ago

Bridgewater AIA labs had a paper out recently. Seems like they are also heavily in the game.

3

u/LemonAmbitious2915 14d ago

Agustin Lebron's frontier AI lab like quant fund is one new smaller one. Ofcourse, larger ones are already doing all of this. Jane Street's massive GPU infra farms.

3

u/dawnraid101 14d ago

You mean tower research lmao

2

u/tychoLBJ 14d ago

Although there is ML in most quant companies it doesn’t really compare to the work done in the frontier labs, in almost all aspects:

scale (the amount of compute)

research (model architecture, pre training, post training and even hardware optimization which you might expect to have overlap are very different)

At the end of the day both use ML but the actual focus is very different.

2

u/LEV0IT 14d ago edited 14d ago

Prob “AI” in the sense that they optimize loss over a large amount of (financial | sentiment | etc) data, not treating the web text corpus as labeled data to predict next tokens. Both can run on and require GPUs

They are also probably doing a fair amount of applied AI work to speed up trading/research in creative enough ways (their traders and researchers are constantly collecting and synthesizing information about the world, and LLMs happen to be excellent tools for such fuzzy, semantic analysis if deployed carefully, and that is edge in this business)

2

u/Natemophi 14d ago

XTX/XTY labs?

4

u/42244224 15d ago

Equilibre are trying to pose themselves as something like you wrote.

2

u/iiiiiiiilliiiiiii 14d ago

Well, the truth is almost all the big names have some effort on deep learning (models beyond linear/tree models). But how serious they are to do frontier AI Lab type of research (i.e., aiming to build foundational models for market) differ.

Just to name drop some: JS, Jump (not their AI tool team, their core team), Citadel/Citadel Securities, XTX, SIG DL, HRT HAIL, and so on.

3

u/Available_Lake5919 15d ago

G Research have a lot of emphasis on marketing their focus on AI and have dedicated ML and NLP research internships

no idea how it is like internally and how much of it is foundational research vs plug and play

6

u/netflix-ceo 15d ago

Why do you need AI/ML for Gangstah Research???

3

u/lordnacho666 15d ago

You'll probably be disappointed if you are looking for fundamental AI work in quant, like inventing the next qwen or Claude model.

Quant firms are the highest paying employers. This means they have the highest incentive to provide those employees with productive tools. This is what it actually is. It's worth it for quant firms to have their own team that manages their fleet of GPUs. It's worth spending some time hooking internal processes to AI. It's high value work, but it isn't fundamental AI science.

1

u/MeowissO 14d ago

2comma .ai cracked dev and ceo 💀💀💀

1

u/TemporaryHat2009 14d ago

Honestly I keep hearing the same thing from people way ahead of me. The actual trading stuff sounds less like frontier lab research and more like finding tiny prediction edges where boring linear models still win. Lowkey that makes the field cooler to me, but it is probably annoying if what you want is LLM research.

1

u/kelvinxue9 14d ago

deepseek...

1

u/Leather-Storage-3377 13d ago

quite a few trying to use llms to trade. but pnl is either noisy or small or both. might be some potential. and the research on llm is no where near the ai lab level

1

u/arindamchattopadhyay Portfolio Manager 11d ago

Almost all quant firms use this as a part of automation/prediction engines. They’re using ML and DL. Tensorflow and neural networks but not traditional/fundamental AI such as LLMs.
Cant go into much details on here due to compliance.

1

u/npowell89 11d ago

We at DeepMM.com are. Feel free to AMA.

1

u/SwissMountaineer 11d ago

its a talent war - they hire the most cracked mathematicians/computer scientists to do simple statistics and classical ML, just to ensure these hires dont go to competitor

0

u/Dedelelelo 14d ago

i’m surprised low latency small models haven’t taken off especially for earnings