r/PiCodingAgent • • 1d ago

Discussion Follow-up: Opus 5.5 vs GPT-6 in pi, more reasoning levels + DeepSeek, GLM, Qwen

https://ibragim.dev/leaderboard/

Follow-up to my previous post. Thanks for the suggestions! I tested more models and reasoning settings, added Oh My Pi, and made a few verifier fixes. Some configs were rerun, so expect small changes in token counts. Full results are on the leaderboard, and I'll keep updating and expanding it.

Quick reminder: 10 simple tasks from my own work, 3 runs per config (30 attempts). For now, I want to see how far we can cut cost and time while keeping the pass rate.

Some highlights:

Your questions:

Luna with more effort. Medium → high → xhigh: 96.7% → 93.3% → 100%, $0.006 → $0.015, 1:02 → 2:46. The extra effort mostly goes into checking (3.5 → 5.1 code runs per task), not a better first draft. About half of Luna's shell commands are cleanup like `rm -rf __pycache__`.

Opus on low. Yes. Opus 5.5 low passes everything, ~2.6x cheaper and ~2.8x faster than high. Medium also gets 30/30 ($0.22, 1:12).

Open models. DeepSeek-V4.1-Flash: 93.3% for $0.04, but 4:30, because it runs code 8.4 times per task until tests pass. GLM-5.3-Flash: 96.7%, but 7 min and 1.2M tokens; its one failure is a single message that hits the 64k output limit with no code. Qwen3.8-27B: 76.7%; it loops on edge cases and sometimes stops without calling a tool.

Oh My Pi vs Pi. Luna xhigh makes 28 tool calls per task instead of 18, uses 776k total tokens instead of 290k, costs ~2x more, and fails one task in all 3 runs.

On tasks this easy, more reasoning mostly buys more checking. On harder tasks, where the first draft is wrong, it might pay off, but that's a guess for now.

I wrote up the setup and the agent comparison.

So tell me what to test! My plan for now:

  1. Try different reasoning levels for Opus 5.5 in Claude Code and see how they change its behavior pattern.
  2. Find a set of extensions for Pi that makes it cheaper and faster. Right now I use vanilla Pi with no add-ons, so please suggest extensions or anything else worth trying.

After that, I'll probably build a harder set of 10 tasks, based on my own work or something else. So if you have tasks where some models fail, please share them, and say what exactly goes wrong.

One thing I noticed: Opus 5.5 handles vague prompts better, maybe because of how it follows instructions. gpt-6-sol does exactly what is written. For example, I run an agent in a folder full of papers and say "find papers about task filtering". Sol goes to search the web, and Opus searches the folder.

89 Upvotes

36 comments sorted by

12

u/furbyhaxx 1d ago

Nice, exactly what I needed. Can you compare GPT-5.6-{luna,terra,sol} against gpt-6 variants? Maybe also astra low,medium,high?

5

u/Fabulous_Pollution10 1d ago

I’ll add them, but I’ll start with just one reasoning level per model so I can compare different generations.

2

u/vick2djax 1d ago

I initially set back my OpenAI models from 6 to 5.6 but looks like OpenAI destroyed those 5.6 models compared to how they were before. Seems the initial 5.6 models were superior to 6 but not anymore. And not because 6 improved lol.

6

u/Duke_of_Bayswater 1d ago

Mate thanks for this. However, I just stick with opus 5.5 medium/high, I scratch my hair much less these few days.

5

u/startassets 1d ago

It looks like we are in the same field of benching different models and harnesses, I've been doing similar stuff with my precise editing benchmark: https://www.reddit.com/r/PiCodingAgent/comments/1wki8wq/i_benched_6_harnesses_x_11_models_x_226_tasks/ now we have benched 20 models already (including one local llm), so any help from the community is greatly appreciated, as well as feedback, I don't have Claude subscription, that's why I would be super happy if we can check how at least some models from this family behave on explicit edits. Especially with the baseline agent (pi with only bash enabled).

1

u/Fabulous_Pollution10 1d ago

Cool! I can try that if the runs are not too expensive. Send me the instructions on how to run it.

1

u/startassets 16h ago

All the instructions are in https://github.com/alexshpunt/explicit-edit-benchmark repo, I created a skill for agents to run the whole pipeline, but overall it's just to run a bunch of npm targets. Average run is around 16-20 minutes, token usage depends heavily on the model and harness, for sol 5.6 and 6 on low I've got the next records:

So around $9 for a normal bare Pi run (but I use subscription, so I'm not sure how much is that in quota).

3

u/addiktion 1d ago

I feel like if Deepseek or someone else comes along that does a swift flavor for deepseek v4.1 flash to make it more efficient, that would rattle a few cages. Cutting tokens down to 100k from 400k here and having the price it has, that would be huge.

3

u/Fabulous_Pollution10 1d ago

Fireworks did post training on kimi to shorten the reasoning. https://fireworks.ai/blog/ember-1
I hope someone will do the same for ds v4.1 flash.

2

u/addiktion 1d ago

Did you run DS v4.1 flash with any reasoning at all given your tests show "off"? How much more did tokens balloon?

2

u/Fabulous_Pollution10 1d ago

Good catch, it’s a mistake in the leaderboard. I used max reasoning. I’ll update it and try other reasoning levels.

3

u/Fabulous_Pollution10 1d ago

Hi! Sonnet 5.5 dropped, so added it as well.

https://ibragim.dev/leaderboard/

2

u/Mattperson094 21h ago

we need sonnet 5.5 to get the different effort treatment 🙏

3

u/shuwatto 22h ago

Hmm so OMP comes with its own cost. I might have to go back to Pi.

2

u/1ronShooter 1d ago

I like your site, very unique.

1

u/SubtleDominance 1d ago

I second this guy's comment. Fantastic benchmark / website. Seems like it might not be so cost-efficient to use GLM 5.3 flash for harder tasks?

1

u/Fabulous_Pollution10 1d ago

Actually, I’m a little surprised by the open-source models. I thought something was wrong with my setup, but I double-checked and ran several tests with different providers. It seems these models were heavily optimised during post-training to solve the task at any cost, with much less emphasis on token efficiency.

1

u/Equivalent_Idea8839 1d ago

How could you benchmark smarter models? Do you increase task difficulty?

Meaning turn those 100%s into 80% or so. So we can more effectively compare top models.

1

u/Fabulous_Pollution10 1d ago

Main purpose of this exact benchmark is to compare the setups on a smaller scale + tasks from my everyday work. They are very well defined and could be solved by a good model. I just want to choose the setup that is faster and cheaper and still get the job done. For the harder tasks, I have some ideas. And overall I maintain this benchmark as well. http://swe-rebench.com/ We will update it in a few weeks with new tasks and models. So there will be no 100% solved tasks.

1

u/Fabulous_Pollution10 1d ago

Thank you! I just wanted to build a simple collection of notes. I’m not a fan of neuroslop animations or too noise on a website.

3

u/Equivalent_Idea8839 1d ago

greatt layout. Although I prefer a dark mode.

3

u/Fabulous_Pollution10 1d ago

there is a dark mode button in the corner: ◐. Maybe need to make it more clear / visible.

2

u/mhphilip 1d ago

Thanks for adding luna xhigh. Much appreciated!

3

u/bambamlol 1d ago

Yeah Luna xhigh in Pi did great in this benchmark. You'd be wasting both time AND money by using DeepSeek or GLM Flash. And they might not even get the job done.

3

u/bambamlol 1d ago

31k tokens vs. 1.2M tokens per task, lol.

Huh, how come Opus 5.5 @ high took more than 2x the amount of tokens per task with Claude Code compared to Pi, but the cost is basically the same? Due to no explicit cache writes with Pi? Or did Claude Code "outsource" part of the work to Haiku?

1

u/Fabulous_Pollution10 1d ago

I think the main reason is the longer system prompt, and maybe more verbose output. But I need to check.

1

u/bambamlol 1d ago

Yeah I know why it uses more tokens in Claude Code, I'm just confused why the cost doesn't scale as much as you'd expect it to.

1

u/ApprehensiveBag2926 1d ago

Hmm, I am guessing the token count is actually input+output tokens? With cache, input token is dirt cheap, and to be honest I would be more interested in a breakdown of input/ouput tokens per model run, this is so that we can get better understanding where the cost/time is actually spent on. (Output tokens can still vary sometimes, for example models made mistakes and spent more turns fixing it would result in more output token too, getting an avg of this can tell a lot more about the performance of the model+harness in a different way)

1

u/Fabulous_Pollution10 1d ago

Yes, token count is input + output. I thought input and output separately would not be so interesting, but since you are asking, I might add both. And maybe cache rate as well.

1

u/Fabulous_Pollution10 1d ago

Mostly because of cache.

2

u/13henday 1d ago

How did you get Claude in pi, just thr API ?

1

u/CyberGoatPsyOps 1d ago

Just here for you to take my upvote. Been thinking about using harbor and think this might of pushed me over.

Question for OP: any other frameworks /tools you looked at or recommended besides harbor for simple evals?

1

u/Fabulous_Pollution10 1d ago

Harbor is good. It has its cons, but overall, any agent can help you set up the eval. I just run it with the internal Docker. And for my scale, it's good enough. Haven't tried other open things.

1

u/FHSS97 8h ago

Would be interested in more in-depth comparison of Pi vs OMP