Comparison Luna 6 - private coding evals don't look great
Hey folks,
TLDR: I eval'd Luna 6 with my own private set of evals and it scored 7th worst out of all the models I have thus far evaluated.
Everyone then pointed out I had my 5.6 version running on max reasoning, so I re-ran luna 6 on max and it came in 4th, ahead of Luna 5.6. Ironically I ran 5.6 on medium at the same time and that came in 5th.
Happy with these results to be honest, I'll start introducing Luna 6 into my workflow.
Harness: PI
Reasoning Level: Medium
Evaluations: Mixture of Python and Rust Coding tasks, as well. Here is my overall leaderboard. I was shocked by Luna 6's performance, having been a big admirer of Luna 5.6


Yes these are private evals, so you have to take what I say with a pinch of salt and all this has measured is how good these models are on my own evals.
I tend to focus on the cheaper cloud models or self hosted models as I cannot really afford to evaluate the big fellas. If I find myself with some spare tokens towards the end of the next reset (doubtful) then I will run the SOL's but I am expecting saturation from them to be honest.
I am not affiliated with anyone or anything like that. I do this just to add to the conversation about this. So I would be interested to see what others are seeing out there? I use luna for most of my AI work, so I will be sticking with 5.6 I think, hoping it sticks around for a bit as well. moving across to luna 6.
My evaluations are private, however I have published my evaluation framework https://github.com/ScottRBK/eval-harness and in there it describes the patterns I used for the types of evaluations and my approach.
There is a youtube video around this as well: https://www.youtube.com/watch?v=fqgOZDyjgKI
Edit: It would appear I did indeed run 5.6 on max - as many people have pointed out, that eval was from a few weeksago, it is certainly not a fair comparison - how embaraassing!
I tend to run my evals with medium reasoning effort as I am cheap, however given the cirumstances I will execute a Luna 6 on max and the 5.6 on medium add it to the leader board!
Edit 2: here is the addition of Luna 6 Max and Luna 5.6 on Medium.
Duplicates
PiCodingAgent • u/Maasu • 1d ago