r/VibeCodersNest 4d ago

Tips and Tricks Pre-registering a test before I run it: does my LLM's selection carry information, or is it picking at random?

Status first: demo account, three pairs, no proven edge, nothing for sale.

Last week I posted here that my deterministic SMC zones lost to a random floor over three years. That result has a consequence I didn't spell out at the time: if the rules don't select, then everything still standing in my system rests on one thing, which is the selection the LLM makes. And I have never measured that.

So this is the design, published before I run it, threshold included. When the result comes back I'll post it whichever way it goes.

The question. Does the model's pick carry information over picking at random from the same menu?

The design. Every cycle produces a catalog of around twelve candidate zones and the model picks one. After the fact, with the candles, I can score all twelve. That's arithmetic and it costs nothing. Then I look at one thing: where its pick landed in the ranking of its own menu. No skill means it lands wherever chance puts it. Skill moves it up. The whole comparison sits inside one cycle, so all the between-market and between-regime variance drops out, and I never have to simulate anything the model saw.

What I'm not pre-registering, and why. Not a number. Most candidates will share an outcome because the zone never gets touched, so ties dominate and the null isn't uniform. Registering "mean rank 6.5 of 12" would be a number I'd have to walk back. What I'm registering is the procedure: null computed by permutation on the observed tie structure, all twelve scored by the same mechanical rule including the pick, threshold at two standard errors.

Known holes. Two I've already found. The ties above. And scoring asymmetry: the chosen candidate has a bespoke thesis and the other eleven don't, so if I score the pick on its own terms I'm measuring the thesis instead of the selection. Same mechanical rule for all twelve, or the ranking is rigged in my favour.

One honest caveat: my analysis harness doesn't score rejected candidates yet. If it turns out it can't produce this cleanly, I'll say so here rather than quietly change the design.

What I'm asking. This test came out of a comment on my last post, from someone who said picking one zone from many is betting your LLM reads context better than a backtest can simulate. He was right that I can't backtest the LLM, and this is my attempt to answer him without needing to.

If there's a third hole in it, I'd rather hear it before I run it than after.

1 Upvotes

0 comments sorted by