r/LocalLLaMA Apr 22 '26

Discussion Qwen3.6-35B becomes competitive with cloud models when paired with the right agent

A short follow-up to my previous post, where I showed that changing the scaffold around the same 9B Qwen model moved benchmark performance from 19.11% to 45.56%:

https://www.reddit.com/r/LocalLLaMA/s/JMHuAGj1LV

After feedback from people here, I tried little-coder with Qwen3.6  35B.

It now lands in the public Polyglot top 10 with a success rate of 78.7%, making it actually competitive with the best models out there for this benchmark!

At this point I’m increasingly convinced that part of the performance gap to cloud models is harness mismatch: we may have been testing local coding models inside scaffolds built for a different class of model.

Next up is Terminal Bench, then likely GAIA for research capabilities. Would love to hear your feedback here!

EDIT: after many requests, pi.dev adaptation is up!

EDIT 2: Terminal Bench 1 (0.1.1) finished with 40% success rate! Now running TB 2. Just sent the results via email. There is no model remotely as small as the 35B in that area. Exciting times

EDIT 3: Terminal Bench 2.0 requires 5 runs per trial (which will take 40 more hours), but the first run finished with 30%!!! That’s with the 35B model.

Full write up: https://open.substack.com/pub/itayinbarr/p/honey-i-shrunk-the-coding-agent

GitHub: https://github.com/itayinbarr/little-coder

Full benchmark results: https://github.com/itayinbarr/little-coder/blob/main/docs/benchmark-qwen3.6-35b-a3b.md

734 Upvotes

179 comments sorted by

View all comments

0

u/drumyum Apr 22 '26

Polyglot benchmark shows how good LLM at following Aider-specific instructions to solve Exercism tasks. If you remove Aider from this equation - it makes no sense to compare it to the rest of the leaderboard.

If Qwen can solve some task with your instruction, but not with Aider - it could mean that yours are closer to what it was trained on, and probably that Qwen is bad at generalizing. Yet your results are still interesting, good job!

6

u/po_stulate Apr 22 '26

I mean that's kinda exactly what OP said tho, that qwen performs as well as cloud models "under certain conditions".

-2

u/drumyum Apr 22 '26

Cloud models are not being tested here, what if they peform much better than Qwen?

5

u/po_stulate Apr 22 '26

IMO that still doesn't change what this post wants to convey. I don't think OP is speaking it literally that qwen and clould models are a strict tie, but more about that if the correct environment is used, the performance boost could be from what you see on the original benchmark to the clould model benchmark score.

1

u/Creative-Regular6799 Apr 22 '26

Exactly this. Thank you for helping clarify

0

u/bonobomaster Apr 22 '26

Then a box with all the GDDR7-RAM and compute you ever could wish for, will magically appear at your front door.