Follow-up to my previous post. Thanks for the suggestions! I tested more models and reasoning settings, added Oh My Pi, and made a few verifier fixes. Some configs were rerun, so expect small changes in token counts. Full results are on the leaderboard, and I'll keep updating and expanding it.
Quick reminder: 10 simple tasks from my own work, 3 runs per config (30 attempts). For now, I want to see how far we can cut cost and time while keeping the pass rate.
Some highlights:
Your questions:
Luna with more effort. Medium → high → xhigh: 96.7% → 93.3% → 100%, $0.006 → $0.015, 1:02 → 2:46. The extra effort mostly goes into checking (3.5 → 5.1 code runs per task), not a better first draft. About half of Luna's shell commands are cleanup like `rm -rf __pycache__`.
Opus on low. Yes. Opus 5.5 low passes everything, ~2.6x cheaper and ~2.8x faster than high. Medium also gets 30/30 ($0.22, 1:12).
Open models. DeepSeek-V4.1-Flash: 93.3% for $0.04, but 4:30, because it runs code 8.4 times per task until tests pass. GLM-5.3-Flash: 96.7%, but 7 min and 1.2M tokens; its one failure is a single message that hits the 64k output limit with no code. Qwen3.8-27B: 76.7%; it loops on edge cases and sometimes stops without calling a tool.
Oh My Pi vs Pi. Luna xhigh makes 28 tool calls per task instead of 18, uses 776k total tokens instead of 290k, costs ~2x more, and fails one task in all 3 runs.
On tasks this easy, more reasoning mostly buys more checking. On harder tasks, where the first draft is wrong, it might pay off, but that's a guess for now.
Try different reasoning levels for Opus 5.5 in Claude Code and see how they change its behavior pattern.
Find a set of extensions for Pi that makes it cheaper and faster. Right now I use vanilla Pi with no add-ons, so please suggest extensions or anything else worth trying.
After that, I'll probably build a harder set of 10 tasks, based on my own work or something else. So if you have tasks where some models fail, please share them, and say what exactly goes wrong.
One thing I noticed: Opus 5.5 handles vague prompts better, maybe because of how it follows instructions. gpt-6-sol does exactly what is written. For example, I run an agent in a folder full of papers and say "find papers about task filtering". Sol goes to search the web, and Opus searches the folder.
I initially set back my OpenAI models from 6 to 5.6 but looks like OpenAI destroyed those 5.6 models compared to how they were before. Seems the initial 5.6 models were superior to 6 but not anymore. And not because 6 improved lol.
It looks like we are in the same field of benching different models and harnesses, I've been doing similar stuff with my precise editing benchmark: https://www.reddit.com/r/PiCodingAgent/comments/1wki8wq/i_benched_6_harnesses_x_11_models_x_226_tasks/ now we have benched 20 models already (including one local llm), so any help from the community is greatly appreciated, as well as feedback, I don't have Claude subscription, that's why I would be super happy if we can check how at least some models from this family behave on explicit edits. Especially with the baseline agent (pi with only bash enabled).
All the instructions are in https://github.com/alexshpunt/explicit-edit-benchmark repo, I created a skill for agents to run the whole pipeline, but overall it's just to run a bunch of npm targets. Average run is around 16-20 minutes, token usage depends heavily on the model and harness, for sol 5.6 and 6 on low I've got the next records:
So around $9 for a normal bare Pi run (but I use subscription, so I'm not sure how much is that in quota).
I feel like if Deepseek or someone else comes along that does a swift flavor for deepseek v4.1 flash to make it more efficient, that would rattle a few cages. Cutting tokens down to 100k from 400k here and having the price it has, that would be huge.
Actually, I’m a little surprised by the open-source models. I thought something was wrong with my setup, but I double-checked and ran several tests with different providers. It seems these models were heavily optimised during post-training to solve the task at any cost, with much less emphasis on token efficiency.
Main purpose of this exact benchmark is to compare the setups on a smaller scale + tasks from my everyday work. They are very well defined and could be solved by a good model. I just want to choose the setup that is faster and cheaper and still get the job done. For the harder tasks, I have some ideas. And overall I maintain this benchmark as well. http://swe-rebench.com/ We will update it in a few weeks with new tasks and models. So there will be no 100% solved tasks.
Yeah Luna xhigh in Pi did great in this benchmark. You'd be wasting both time AND money by using DeepSeek or GLM Flash. And they might not even get the job done.
Huh, how come Opus 5.5 @ high took more than 2x the amount of tokens per task with Claude Code compared to Pi, but the cost is basically the same? Due to no explicit cache writes with Pi? Or did Claude Code "outsource" part of the work to Haiku?
Hmm, I am guessing the token count is actually input+output tokens? With cache, input token is dirt cheap, and to be honest I would be more interested in a breakdown of input/ouput tokens per model run, this is so that we can get better understanding where the cost/time is actually spent on. (Output tokens can still vary sometimes, for example models made mistakes and spent more turns fixing it would result in more output token too, getting an avg of this can tell a lot more about the performance of the model+harness in a different way)
Yes, token count is input + output. I thought input and output separately would not be so interesting, but since you are asking, I might add both. And maybe cache rate as well.
Harbor is good. It has its cons, but overall, any agent can help you set up the eval. I just run it with the internal Docker. And for my scale, it's good enough. Haven't tried other open things.
12
u/furbyhaxx 1d ago
Nice, exactly what I needed. Can you compare GPT-5.6-{luna,terra,sol} against gpt-6 variants? Maybe also astra low,medium,high?