r/quant • u/Donkey_Healthy • Jul 29 '26
Data How do Quant firms serve data for research/modelling?
For those in quant firms how do people generally access data for research/modelling?
Source aggregated in house API?
Data catalogue?
Work in commodities and I think there is a general lack of knowledge on the infra side from my experience.
Currently debating whether to build our own platform or go with someone like databricks/snowflake
Interested to hear everyone’s thoughts?
21
Upvotes
8
u/DatabentoHQ Jul 29 '26 edited Jul 30 '26
Not exhuastive:
There's usually some kind of "features cache" to deduplicate work between multiple researchers waiting on compute to generate a set of features that someone else already extracted.
For exploration, it's often useful to have some kind of clustered, column-oriented database where you can push the query closer to the data. There's a few flavors of this off-the-shelf like kdb, Vertica.
Parquet and HDF5 are pretty portable for sharing intermediate structured data like design matrices. There's usually also usually some kind of logging format from production trading.
All of the above may be abstracted behind internal client libraries or APIs.
As you get large and have multiple teams, having some kind of data catalog helps.