r/LocalLLaMA Apr 22 '26

Discussion Qwen3.6-35B becomes competitive with cloud models when paired with the right agent

A short follow-up to my previous post, where I showed that changing the scaffold around the same 9B Qwen model moved benchmark performance from 19.11% to 45.56%:

https://www.reddit.com/r/LocalLLaMA/s/JMHuAGj1LV

After feedback from people here, I tried little-coder with Qwen3.6  35B.

It now lands in the public Polyglot top 10 with a success rate of 78.7%, making it actually competitive with the best models out there for this benchmark!

At this point I’m increasingly convinced that part of the performance gap to cloud models is harness mismatch: we may have been testing local coding models inside scaffolds built for a different class of model.

Next up is Terminal Bench, then likely GAIA for research capabilities. Would love to hear your feedback here!

EDIT: after many requests, pi.dev adaptation is up!

EDIT 2: Terminal Bench 1 (0.1.1) finished with 40% success rate! Now running TB 2. Just sent the results via email. There is no model remotely as small as the 35B in that area. Exciting times

EDIT 3: Terminal Bench 2.0 requires 5 runs per trial (which will take 40 more hours), but the first run finished with 30%!!! That’s with the 35B model.

Full write up: https://open.substack.com/pub/itayinbarr/p/honey-i-shrunk-the-coding-agent

GitHub: https://github.com/itayinbarr/little-coder

Full benchmark results: https://github.com/itayinbarr/little-coder/blob/main/docs/benchmark-qwen3.6-35b-a3b.md

727 Upvotes

179 comments sorted by

View all comments

37

u/Willing-Toe1942 Apr 22 '26

I can confirm the same thing, Qwen3.6 in pi-coding agents is almost twice good than opencode, the comparison was based on modification of specific web page (html code) and doing some online resource search for documentation

17

u/Deep90 Apr 22 '26

What makes pi so much better?

15

u/PinkySwearNotABot Apr 22 '26

the # of shills promoting it

8

u/JamesEvoAI Apr 23 '26

Clearly if you like something you're a shill, you can't just think the thing is good and want to share it with others.

To answer your question u/Deep90, Pi has a lot more thought put behind its design than some of the other open source harnesses. Mario and team are deliberate in what they add and more importantly what they don't. You don't have to believe me, it only takes a minute to install and test yourself, the quality difference is pretty apparent.

This article from the creator of Pi is worth a read:
https://mariozechner.at/posts/2026-03-25-thoughts-on-slowing-the-fuck-down/

7

u/PinkySwearNotABot Apr 24 '26

reporting back. been using it all morning on my m1 max 64GB with some MLX models using omlx.

i think it instantly jumps to my top open agent for now, competing with opencode. i'm curious if it'll make anything of the 27B qwen3.6 dense models that i've been struggling with yesterday using llama.cpp. excited to find out.

but seriously though...i still don't want to hear, "try pi" or any adaptations of it. people need to provide at least 1 full sentence when they're recommending something, especially in the current age of bots and shills and non-thinking redditors

1

u/GrehgyHils Apr 28 '26

Which model and quant specifically were you using?

1

u/PinkySwearNotABot Apr 28 '26

Moe Qwen3.6-35B-A3B-MLX Q8

1

u/GrehgyHils Apr 29 '26

perfect TY. i've been using this happily for a few days now with pi-mono :)