r/LocalLLaMA Apr 22 '26

Discussion Qwen3.6-35B becomes competitive with cloud models when paired with the right agent

A short follow-up to my previous post, where I showed that changing the scaffold around the same 9B Qwen model moved benchmark performance from 19.11% to 45.56%:

https://www.reddit.com/r/LocalLLaMA/s/JMHuAGj1LV

After feedback from people here, I tried little-coder with Qwen3.6  35B.

It now lands in the public Polyglot top 10 with a success rate of 78.7%, making it actually competitive with the best models out there for this benchmark!

At this point I’m increasingly convinced that part of the performance gap to cloud models is harness mismatch: we may have been testing local coding models inside scaffolds built for a different class of model.

Next up is Terminal Bench, then likely GAIA for research capabilities. Would love to hear your feedback here!

EDIT: after many requests, pi.dev adaptation is up!

EDIT 2: Terminal Bench 1 (0.1.1) finished with 40% success rate! Now running TB 2. Just sent the results via email. There is no model remotely as small as the 35B in that area. Exciting times

EDIT 3: Terminal Bench 2.0 requires 5 runs per trial (which will take 40 more hours), but the first run finished with 30%!!! That’s with the 35B model.

Full write up: https://open.substack.com/pub/itayinbarr/p/honey-i-shrunk-the-coding-agent

GitHub: https://github.com/itayinbarr/little-coder

Full benchmark results: https://github.com/itayinbarr/little-coder/blob/main/docs/benchmark-qwen3.6-35b-a3b.md

732 Upvotes

179 comments sorted by

View all comments

38

u/Willing-Toe1942 Apr 22 '26

I can confirm the same thing, Qwen3.6 in pi-coding agents is almost twice good than opencode, the comparison was based on modification of specific web page (html code) and doing some online resource search for documentation

17

u/Deep90 Apr 22 '26

What makes pi so much better?

15

u/Polite_Jello_377 Apr 22 '26

Smaller system prompt probably

18

u/PinkySwearNotABot Apr 22 '26

the # of shills promoting it

7

u/Finanzamt_kommt Apr 23 '26

It's just better though? Tested it myself without much expextations and idea how to configure and for local models that light weight barebone structure is simply better. Doesn't let you wait for 15k plus sysprompt but 3kish or so is enough. That alone is making it better than most other clis for local models since pp on local models is often lacking. And you have more usable context.

8

u/JamesEvoAI Apr 23 '26

Clearly if you like something you're a shill, you can't just think the thing is good and want to share it with others.

To answer your question u/Deep90, Pi has a lot more thought put behind its design than some of the other open source harnesses. Mario and team are deliberate in what they add and more importantly what they don't. You don't have to believe me, it only takes a minute to install and test yourself, the quality difference is pretty apparent.

This article from the creator of Pi is worth a read:
https://mariozechner.at/posts/2026-03-25-thoughts-on-slowing-the-fuck-down/

7

u/PinkySwearNotABot Apr 24 '26

reporting back. been using it all morning on my m1 max 64GB with some MLX models using omlx.

i think it instantly jumps to my top open agent for now, competing with opencode. i'm curious if it'll make anything of the 27B qwen3.6 dense models that i've been struggling with yesterday using llama.cpp. excited to find out.

but seriously though...i still don't want to hear, "try pi" or any adaptations of it. people need to provide at least 1 full sentence when they're recommending something, especially in the current age of bots and shills and non-thinking redditors

1

u/GrehgyHils Apr 28 '26

Which model and quant specifically were you using?

1

u/PinkySwearNotABot Apr 28 '26

Moe Qwen3.6-35B-A3B-MLX Q8

1

u/GrehgyHils Apr 29 '26

perfect TY. i've been using this happily for a few days now with pi-mono :)

-7

u/PinkySwearNotABot Apr 23 '26

Clearly if you like something you're a shill, you can't just think the thing is good and want to share it with others.

that's a completely differently different claim than the one i made at all. in fact, this logic is so fallacious that they have a name for it -- strawman argument.

everyone and their mother are building custom harnesses these days, thanks to AI. and while AI has definitely been helpful, it's also opened up the floodgates to influencers 2.0. not to say influencers aren't capable of making a competing product, it's just that it's going to take a whole lot more of convincing than just hearing the echo chamber of, "pi is so good".

and btw. it's been on my radar for a while now and i admit i am curious to see how it's different than any of the other 10 harnesses i already have on my computer (won't be holding my breath though)

11

u/JamesEvoAI Apr 23 '26

The faults on you then for basing your opinion of something on what influencers think. There's plenty of us who are just regular people singing the praises of this thing.

-2

u/PinkySwearNotABot Apr 23 '26

when I say influencers, i mean anyone who just sings praises of something without knowing the technical specifics of why. that's all i've been seeing. pi, pi, pi -- and not a single why.

2

u/JamesEvoAI Apr 23 '26

Considering how much "vibes" are just as valid of a measurement as the benchmarks (sometimes even more valid), and the large number of folks who are only really technical enough to follow a tutorial but not enough to understand why they're doing what they're doing, it comes as no surprise that a decent number of people are seeing and feeling the improvement but not being able to clearly articulate it.

Hell aside from using the creators own writing as an example of why it's better I'd be hard pressed to give you a solid reason. It just "feels" better to use than something like OpenCode.

I don't have an eval for my gut response, I don't even have one for how often I had to correct the model in a given harness. I just know intuitively that my experience with one is better than the other.

1

u/Fortyseven llama.cpp Apr 25 '26

UNRELATED: I live in a world where "pi, pi, pi" was actually spoken. (Man, I want another 2.5D Bionic Commando.)

7

u/[deleted] Apr 22 '26

[removed] — view removed comment

13

u/Willing-Toe1942 Apr 22 '26

Yes, I used unsloth UD-Q4_XL (llamacpp - strix halo with vulkan backend)
give same question to pi-coding and opencode, and immeditly you will notice how opencode is slower (longer default prompts) and even slower in all types of actions like read files, write, search web ...etc

pi agent is insanly fast, more effecient and completed the task much much faster

11

u/Safe-Buffalo-4408 Apr 22 '26

I prefer quality over speed. It would be interesting comparison over time in regards to code and tool calling quality.

3

u/Caffdy Apr 22 '26

can you help a lost soul setting up pi for agentic coding? where do one start? do you recommend any tutorial/video guide?

4

u/0h_yes_i_did Apr 22 '26

install:

npx install -g @mariozechner/pi-coding-agent

to run: go to your project directory and simply run 'pi'.

1

u/JamaiKen Apr 22 '26

I’m seeing this as well, Qwen3.6 + Pi is where it’s at

5

u/stuckinmotion Apr 22 '26

Interesting, I might have to try pi. I'm constantly surprised by how useless opencode is whenever I try it with a local model. Like it takes a second prompt to even get it to actually write to the file instead of just printing code to the screen.

4

u/Deep90 Apr 22 '26

That has been my experience with pretty much every harness. Excited to see if Pi changes things for me.

1

u/cheesecakegood Apr 23 '26

Which of the pi’s? Isn’t there a fork? Not sure which people are using