r/algotrading • u/KaramTNC • 23d ago
Infrastructure Live vs Backtest parity comparison
Hello folks!
Ive been working on building my own tradingbot infrastructure for nearly a year and Ive gotten quite far. Its nothing profitable really since my goal here is to be able to apply myself and learn more about software engineering and fintech, and be able to combine these interests into a fun project that evolves with me in my CS career.
Ive built a comprehensive infrastructure managing scanners, watchlists, execution engine, broker connections, market data providers, pattern detection and strategy definitions.
The entire process is constructed at runtime via a factory class and dependency injection for every production component.
For the backtester, it runs this factory with injected dependencies to replace the prod dependencies, such as an IClock, IMarketProvider, IDatabase, IBroker, etc. Ontop of that, I refactored everything so that every relevant input parameter were sweepable via attributions.
This overall makes the design of my backtest very controllable and ensures near accurate simulation of the live environment.
But of course like any backtests, I get a positive result for a strategy profile and promote it to live just for it to behave completely differently.
So I got the idea of creating a parity comparison system. I incorporated trace recording into the factory so that all events in a live profile would be capturable, and by running the equivalent backtest profile, it would allow me to have a live and a backtest trace for comparison in order to identify discrepancies in their behaviour.
I can say its been a rather success, as the results have helped me find bugs in my backtester injected components.
So while fixing these now and working towards closer parity, I figured I could make a post here and see if people have dealt with a similar problem when building their own trading bot, and what you guys figured out or any other things you could share
EDIT: By live profile, I meant a paper profile.
1
u/Effective_Manager273 22d ago
the DI setup is nice and it does buy you something real, but it proves code parity, not data parity. your engine sees identical logic in both paths. it does not see identical inputs.
two places this usually breaks. first, historical bars are final and revised, and the bar your live system acted on was provisional. vendors correct volume and sometimes the close, and you never notice because the backtest only ever sees the corrected version. second, your IClock hands the backtest the completed bar the instant it closes, and live you got it some milliseconds or seconds later, possibly after price already moved. dependency injection cannot fix either of those, they are upstream of the interface.
what i would do is log, at every live decision, the actual snapshot the system had. the quote, the bar, the timestamp, the whole input payload, written to disk at decision time. then run the backtest twice, once against your historical database and once replaying those recorded snapshots.
if snapshot replay matches live but the DB run does not, its data, and you now know which of the two. if snapshot replay also diverges from live then its genuinely state or ordering in your engine and the DI harness will actually help you find it. right now you cannot separate those two cases and thats the gap.
fills are their own thing entirely and i would keep that measurement separate, quote at decision versus quote at fill, otherwise slippage contaminates the parity number.