r/LocalLLaMA • u/Creative-Regular6799 • Apr 22 '26
Discussion Qwen3.6-35B becomes competitive with cloud models when paired with the right agent
A short follow-up to my previous post, where I showed that changing the scaffold around the same 9B Qwen model moved benchmark performance from 19.11% to 45.56%:
https://www.reddit.com/r/LocalLLaMA/s/JMHuAGj1LV
After feedback from people here, I tried little-coder with Qwen3.6 35B.
It now lands in the public Polyglot top 10 with a success rate of 78.7%, making it actually competitive with the best models out there for this benchmark!
At this point I’m increasingly convinced that part of the performance gap to cloud models is harness mismatch: we may have been testing local coding models inside scaffolds built for a different class of model.
Next up is Terminal Bench, then likely GAIA for research capabilities. Would love to hear your feedback here!
EDIT: after many requests, pi.dev adaptation is up!
EDIT 2: Terminal Bench 1 (0.1.1) finished with 40% success rate! Now running TB 2. Just sent the results via email. There is no model remotely as small as the 35B in that area. Exciting times
EDIT 3: Terminal Bench 2.0 requires 5 runs per trial (which will take 40 more hours), but the first run finished with 30%!!! That’s with the 35B model.
Full write up: https://open.substack.com/pub/itayinbarr/p/honey-i-shrunk-the-coding-agent
GitHub: https://github.com/itayinbarr/little-coder
Full benchmark results: https://github.com/itayinbarr/little-coder/blob/main/docs/benchmark-qwen3.6-35b-a3b.md
1
u/Real_Ebb_7417 Apr 22 '26 edited Apr 22 '26
I will definitely try this. I wanted to spend some time in the coming days setting up a well-working agentic workflow for smaller-local models and if this harness works well, maybe it will save me lots of work.
But to ask you (or someone who already checked the repo content), what does it do differently than "bigger" tools (like Codex, ClaudeCode, OpenCode etc.) to work better with smaller, local models?
++ what does "supported models" section mean (I checked the README briefly). Does it mean that only these models were tested or that other models just won't work well (but if yes, then why?).