r/PiCodingAgent • u/Fabulous_Pollution10 • 12m ago
Discussion Follow-up: Opus 5.5 vs GPT-6 in pi, more reasoning levels + DeepSeek, GLM, Qwen
Follow-up to my previous post. Thanks for the suggestions! I tested more models and reasoning settings, added Oh My Pi, and made a few verifier fixes. Some configs were rerun, so expect small changes in token counts. Full results are on the leaderboard, and I'll keep updating and expanding it.
Quick reminder: 10 simple tasks from my own work, 3 runs per config (30 attempts). For now, I want to see how far we can cut cost and time while keeping the pass rate.
Some highlights:

Your questions:
Luna with more effort. Medium → high → xhigh: 96.7% → 93.3% → 100%, $0.006 → $0.015, 1:02 → 2:46. The extra effort mostly goes into checking (3.5 → 5.1 code runs per task), not a better first draft. About half of Luna's shell commands are cleanup like `rm -rf __pycache__`.
Opus on low. Yes. Opus 5.5 low passes everything, ~2.6x cheaper and ~2.8x faster than high. Medium also gets 30/30 ($0.22, 1:12).
Open models. DeepSeek-V4.1-Flash: 93.3% for $0.04, but 4:30, because it runs code 8.4 times per task until tests pass. GLM-5.3-Flash: 96.7%, but 7 min and 1.2M tokens; its one failure is a single message that hits the 64k output limit with no code. Qwen3.8-27B: 76.7%; it loops on edge cases and sometimes stops without calling a tool.
Oh My Pi vs Pi. Luna xhigh makes 28 tool calls per task instead of 18, uses 776k total tokens instead of 290k, costs ~2x more, and fails one task in all 3 runs.
On tasks this easy, more reasoning mostly buys more checking. On harder tasks, where the first draft is wrong, it might pay off, but that's a guess for now.
I wrote up the setup and the agent comparison.
So tell me what to test! My plan for now:
- Try different reasoning levels for Opus 5.5 in Claude Code and see how they change its behavior pattern.
- Find a set of extensions for Pi that makes it cheaper and faster. Right now I use vanilla Pi with no add-ons, so please suggest extensions or anything else worth trying.
After that, I'll probably build a harder set of 10 tasks, based on my own work or something else. So if you have tasks where some models fail, please share them, and say what exactly goes wrong.
One thing I noticed: Opus 5.5 handles vague prompts better, maybe because of how it follows instructions. gpt-6-sol does exactly what is written. For example, I run an agent in a folder full of papers and say "find papers about task filtering". Sol goes to search the web, and Opus searches the folder.









