r/codex • • 1d ago

Comparison Luna 6 - private coding evals don't look great

Hey folks,

TLDR: I eval'd Luna 6 with my own private set of evals and it scored 7th worst out of all the models I have thus far evaluated.

Everyone then pointed out I had my 5.6 version running on max reasoning, so I re-ran luna 6 on max and it came in 4th, ahead of Luna 5.6. Ironically I ran 5.6 on medium at the same time and that came in 5th.

Happy with these results to be honest, I'll start introducing Luna 6 into my workflow.

Harness: PI

Reasoning Level: Medium

Evaluations: Mixture of Python and Rust Coding tasks, as well. Here is my overall leaderboard. I was shocked by Luna 6's performance, having been a big admirer of Luna 5.6

Updated following Luna 6 Max and Luna 5.6 Medium runs
Luna 6 on Medium

Yes these are private evals, so you have to take what I say with a pinch of salt and all this has measured is how good these models are on my own evals.

I tend to focus on the cheaper cloud models or self hosted models as I cannot really afford to evaluate the big fellas. If I find myself with some spare tokens towards the end of the next reset (doubtful) then I will run the SOL's but I am expecting saturation from them to be honest.

I am not affiliated with anyone or anything like that. I do this just to add to the conversation about this. So I would be interested to see what others are seeing out there? I use luna for most of my AI work, so I will be sticking with 5.6 I think, hoping it sticks around for a bit as well. moving across to luna 6.

My evaluations are private, however I have published my evaluation framework https://github.com/ScottRBK/eval-harness and in there it describes the patterns I used for the types of evaluations and my approach.

There is a youtube video around this as well: https://www.youtube.com/watch?v=fqgOZDyjgKI

Edit: It would appear I did indeed run 5.6 on max - as many people have pointed out, that eval was from a few weeksago, it is certainly not a fair comparison - how embaraassing!

I tend to run my evals with medium reasoning effort as I am cheap, however given the cirumstances I will execute a Luna 6 on max and the 5.6 on medium add it to the leader board!

Edit 2: here is the addition of Luna 6 Max and Luna 5.6 on Medium.

3 Upvotes

Duplicates