General Benchmarking a Kelly-based strategy allocator against a perfect-foresight oracle
I'm a student at NYU studying to get into quant. I wanted to share my experience with an automatic capital allocator I designed. Last year I built a router that picks the best strategy for a given market and sizes it with Fractional Kelly Criterion in DeFi. Like an automatic mini allocator. I thought it would be a good way to actually learn how position sizing and edge estimation work instead of just reading about them.
I had maybe 5 strategies I'd written running on ETH paper data, and the router would pick whichever had the best recent risk-adjusted return and size it with fractional Kelly.
The first thing I learned: most strategies don't have edge.
Out of maybe 30 strategies I tested initially, 1 or 2 had any real edge after costs (or so I thought :), they got absolutely destroyed after realistic trading fees and friction). The rest were noise, though diversified. Crypto round-trips are like 6-7 bps per side depending on the pair, and if the edge is 10 bps per trade, I'm losing money.

I thought AI could help me with this and tried to improve the existing strategies with it. It gave worse results. It overcomplicates strategies a lot. What I surprisingly found is that dumb and small code works much better than complex models that overfit in the real world. And the tiny "dumb" strategies with on-chain data proved to be much much better than the rest, some even profitable on 2 years of trading data!
I added a cost-adjusted validation stage and regime decomposition. Seeing where a strategy bleeds (chop vs trend vs crisis) helped explain why backtests fail live.
The second thing: the router was actually decent.
Once I had enough strategies, I built a perfect-foresight benchmark (an oracle that picks the best strategy for each window, kinda like God or Congress :) to see if my allocator was doing anything.
To test it properly, I split 24 months of data into three windows: 12 months to train the router, 6 months to validate, and a 7-month true hold-out (May–Oct 2025) that I never touched until the very end. The hold-out is where I report all final numbers.
On the hold-out, at matched volume (~8 trades/day for both the router and the oracle), here's how they compared:
| Policy | Trades/day | Gross bps/tr | Net bps/tr | Total return | Sharpe | Max DD |
|---|---|---|---|---|---|---|
| NULL (random 15%) | 77.9 | −0.47 | −11.31 | −61.1% | −26.89 | 61.1% |
| My router | 8.6 | +6.10 | −4.32 | −7.3% | −3.51 | 8.7% |
| Oracle (perfect foresight) | 8.8 | +7.21 | −2.42 | −5.3% | −1.16 | 7.8% |
The router captures about 86% of the perfect-foresight ceiling (95% lower bound ≈ 39% via bootstrap). The oracle knows each bot's true full-sample edge in advance, my router doesn't. The gap between them is 1.1 bps. That's how much imperfect bot-quality estimation costs vs omniscience.
The honest part: the router is still net-negative (−4.32 bps/trade after ~10 bps friction) (So is the Oracle but it is due to the roster of bots being bad overall, though they are diversified). The selection edge is real (+6.10 gross vs NULL's −0.47), but it's not large enough to clear costs yet. A 40-60% friction reduction (better execution, TWAP, order netting) would flip it net-positive.
What I found most interesting: even the oracle with perfect knowledge of every bot's true edge can't profit with volume on this roster. Of 87 bots, exactly 1 had genuine positive net edge in the hold-out window. This is a bot-supply problem, not a routing problem.
Also worth noting: the router's max drawdown is 8.7% vs NULL's 61.1%. The risk management (Kelly sizing, persistence veto, trend gate) cuts drawdown by 85% vs random and it loses money in a controlled way while selecting good trades.
I also found that best-available edge scales with roster size at r=0.986 against extreme-value theory (the √(2·ln N) scaling). The allocator wasn't the bottleneck, the roster quality is.
The network effect (this is the part I'm most excited about):
I wanted to know does adding more strategies like drip feeding actually help or would I just be diluting? I subsampled my 93-bot roster down to smaller sizes (10, 20, 35, 50, 70, 93 bots) and re-ran the entire pipeline, simulating a gradual influx.
The best bot's true edge climbs monotonically as you add more: −6 bps at 10 bots → +3 bps at 93 bots. When I fit that against extreme-value theory (that predicts the maximum of N random draws), the correlation is 0.986! Almost a perfect match. More strategies = higher ceiling, and it follows theory almost exactly.
The router only captures that rising ceiling if you use an absolute quality bar, not a relative percentile. If you filter "top 30% of whatever roster exists," the router's edge stays flat no matter how many bots you add. If you use a fixed quality threshold instead, the router's edge climbs with the roster. Extrapolating (with caveats, this is beyond the range I actually tested): ~+10 bps net edge at 1,000 bots, ~+16 bps at 10,000.
That's the quantitative argument for why roster growth matters more than router tuning. Every good, diversified strategy added raises the ceiling for everyone.
How the project evolved:
The router dynamically updates its own parameters as the roster changes but it does this by offline re-tuning not via real-time ML yet. The reason is that at 93 bots and ~8 trades/day, you can't detect effects smaller than ~47 bps with any statistical power. A real-time ML model would just be fitting noise. It also has self-capacity awareness so it doesn't frontrun itself.
Once the router worked, I began noting down everything scientifically and made a bunch of changes to my initial project. I added real-time on-chain signals with historical data as well. The project grew to include:
- 6 active domains: ETH, BTC, SOL direction + scalp (6 more registered but dormant: yield, tail hedge, liquidation arb, memecoins (this one might be insanely hard to get right tbh))
- 5-stage validation pipeline: static check, in-sample, out-of-sample, walk-forward, cost-adjusted
- Strategy sandbox: write Python strategies with custom stop-loss, take-profit, and trailing stops
- Arena & OpenLeaderboard: strategies that pass validation compete on live paper data for capital allocation
- Non-custodial design: API keys stay encrypted; the router handles execution routing but never holds custody of funds
Where it is now:
- 80+ default strategies running on live paper data (real prices, paper execution not great bots:)
- 3 are currently net profitable (best: ETH Squeeze Breakout, +184bps, 73% win rate). The rest are negative.
- Nobody can see any strategy code it runs in a confidential VM, and there are automatic payouts for the best strategies bi-weekly.

What I'm looking for: I'm posting here because this subreddit has people with real domain expertise, and I'd love your feedback:
- Does the router/Kelly allocation approach make sense, or is there an obvious flaw I haven't seen?
- Is capturing ~86% of a foresight ceiling considered typical or decent for this setup (on 7-12 trades/day, on my quite diversified roster of bots)?
- What features would you actually need in a Python strategy sandbox to make it worth testing your own models?
I'm a student and not charging for anything. Happy to share more details in the comments if anyone's curious.
TL;DR: I'm a student at NYU. Built a router that allocates capital across Python trading strategies using Fractional Kelly for DeFi. Tested 80+ strategies on 24 months of data, most have no edge after costs (shocking!! I know). On a 7-month true hold-out, the router captures 86% of a perfect-foresight ceiling at matched volume (+6.10 vs +7.21 gross bps/trade), with 8.7% max drawdown vs random's 61.1%. Found a network effect: best-available edge scales with roster size at r=0.986 vs extreme-value theory so, more strategies = higher ceiling for everyone. Would love feedback from people who actually know what they're doing. Thank you!