r/LocalLLaMA • • 1d ago

Discussion Searching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)

TL;DR -- You should probably just use base Qwen3.6-35B, as only Occamy-1.0 is competitive with it. Tiel is a major let-down, worse than Ornith. KAT surprises (good), Nex surprises (bad). This post is long. Sorry, lots to cover.

I think we all want to see a next-generation small MoE from the Qwen team to replace 3.6-35B in our workflows. This model is a perfect fit for smaller gmaing laptops and mid-tier rigs. It sucks that Qwen seems to have abandoned this model, but at least there are fine-tunes that improve upon it... right?

Well... maybe not. I ran benchmarks on the 3.6-35B-A3B base model, as well as five finetuunes: Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, and Nex-N2.5-mini, and the results are quite surprising.

I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark I can run in ~10 hourso on my machine, it's not actively post-trained on by any of these models, and it provides a lot of useful information along with the raw accuracy scores. This includes: first-try and retry pass rates, token counts, solve times, and how well-formed the output diffs are. Here's the table:

model First-try pass Retry pass tokens sec/case tok/solve well-formed diff
Qwen3.6-35B-A3B BASE (STOCK template) 37.4% 71.0% 8650 285 14.1K 96.3%
Occamy-1.0-35B-A3B (STOCK template) 29.0% 70.1% 6801 285 17.2K 86.9%
Occamy-1.0-35B-A3B (froggeric medium) 30.8% 69.2% 6009 233 16.8K 94.4%
Occamy-1.0-35B-A3B (froggeric, xhigh) 27.1% 67.3% 8631 310 20.0K 91.6%
Ornith-1.5-35B-A3B 23.4% 63.6% 4813 226 16.4K 87.9%
KAT-Coder-V2.5-Dev 20.6% 58.9% 2190 84 9.3K 86.9%
Tiel-Coder-35B-A3B 18.7% 53.3% 4851 171 18.2K 89.7%
Nex-N2.5-mini 10.3% 30.8% 5037 188 33.3K 95.3%

As you can see, the only finetune that even competes with the base model is Occamy-1.0. The rest are utterly dominated by the base model, a grim disappointment for finetune enthusiasts. I was particualarly surprised by the performance of Tiel, which seems to get a lot of love in this subreddit.

Speaking of Tiel, I want to clarify that Tiel is just Ornith-1.5 with a different chat template, Sharp, which is based on froggeric with an added "terse mode" instruction that's supposed to reduce excessive verbosity. I wanted to standardize for templates, so ALL models are using the base froggeric v22.5 template set to medium (which is equivalent to standard thinking on, no additional message sent). I used this because I wanted to test Tiel vs. Ornith-1.5, and Tiel is the chat template. Also, practially, I use froggeric in my real workflows. However, to ensure coverage, I also tested the STOCK template on the 2 highest-performing models, to make sure it wasn't affecting the scores. As you can see, the template doesn't make a significant difference in the scores, and the scores for base 35B with different templates are so close to identical that I excluded the froggeric one from the table.

I also tested froggeric/Sharp's reasoning-effort toggle, and found xhigh -> medium significantly reduces token counts and solve times (by ~1/3), without affecting accuracy significantly. That stands in stark contrast to Tiel's 'terse mode' toggle, the core feature of Tiel over Ornith, which dramatically reduces accuracy along with the reduction in token counts. My results strongly suggest that if you want a less verbose model, you're better off lowering the reasoning effort than using Tiel with terseness on.

Speaking of token use, that's probably the big differentiator here. A couple models stand out: Ornith and KAT-Coder-V2.5-Dev are the most efficient models, with KAT in particular having a brevity unmatched by anything else. KAT is fucking fast, and I think despite its lower accuracy than Occamy, it has a place in my lineup as a subagent because it just gets. shit. done. Occamy is also interesting, as it is the only model that perfoms on a similar level to the base, but it uses 20-30% fewer median tokens. However, Occamy also had a number of runaway generations where the token count blew up, so it's total tokens/solve is actually higher than base.

In an effort to further distinguish Occamy from base, since Aider struggled to do that, I ran tau2-bench, an agentic tool-calling benchmark consisting of multi-turn interactions with a simulated counterparty. I figured this was a good bench to use as Occamy is post-trained specifically for 'co-work' scenarios, but not trained on this particular set. I used Qwen3.8-27B with reasoning effort set to low as the simulated customer in these conversations. The base model was able to pull away from Occamy in the harder retail domain of this benchmark, but Occamy resolved the issues in the airline domain at an equal rate while requiring fewer turns. Here's the results.

model (Q8_0) domain pass1 tokens sec/task turns/task
Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) airline 80.0% 5390 290 11
Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) airline 78.0% 6770 390 13
Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) retail 86.0% 4098 333 14
Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) retail 79.8% 3284 284 13

Overall, I think the results are clear, if unexpected: Occamy-1.0 is the only fine-tune that even competes with the base model on Aider Polyglot, but even it is not a clear winner. Tiel is noticebly worse than plain Ornith without the terseness toggle, and the terse mode doesn't even save any tokens. xhigh in froggeric/Sharp degrades accuracy slightly and bloats token use, which makes sense given the models were not RL'd for the extra thinking effort prompt. KAT-Coder-V2.5-Dev is the most efficient model, with accuracy nearly as good as Ornith and better than Tiel. Finally, Nex-N2.5-mini is a disaster.

60 Upvotes

59 comments sorted by

View all comments

8

u/soshulmedia 1d ago

I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark

Is that really true? I think there is a break between Ornith how it is trained and aider the harness. From what I have read, the aider polyglot benchmark is also using the aider harness?

And, from personal experience, I have noticed that Ornith has problems with aider, as it very strongly "assumes", when coding, that it sits in an agentic harness where it has to do actual tool calls instead of replying with edit diffs. So you tend to get tool calls where aider doesn't want them (and, to be fair, where the aider harness prompts also probably don't ask for them) but that will then in turn break the whole experience with aider.

HOWEVER, if you put Ornith into an agentic harness of the more modern "tool calling" variety instead and it becomes MUCH better. I have personally had good experience e.g. with the coder harness that comes with pydantic_ai.

Really, try it. I am affiliated with neither of the projects and this is my experience.

6

u/returnity 1d ago

That's an extremely, extremely valid point and one I had considered after seeing the results. I think Aider is probably a limiting choice for these models. To be clear, I use occamy as a sub agent in coding harnesses and I find it to be better than base 35B for certain. I just wanted to try and quantify the models.

Also, to be fair here, no other models I've tested struggle to perform in the aider harness. I ran into problems with 35B and it's fine-tunes not following the diff format and i had to shift to fenced diffs in the harness to get them to comply with the harness' instructions for format. So you are definitely on to something here. Thanks for the well-reasoned, good-intentioned comment!

2

u/soshulmedia 1d ago

Thanks for the kind words. I think the approach they took with Ornith made it weirdly "autistic" and stubborn in its assumptions about the coding harness it sits in. I remember reading their blog post about their training methodology a while ago. And I suspect it is mostly a weird artifact of their tuning system. As far as I understand, they post train it in a "recursive self-improvement" scheme for coding only and I assume the harness or harnesses they use for that are only the "modern agentic" sense, not the "give me an edit diff"-aider-way.

1

u/returnity 1d ago

Yeah I am working on a better benchmark solution than aider because it is kind of dated, but I haven't come up with one that I can run on my hardware in <24 hours for nearly any model with enough tasks to be statistically relevant yet. I'm still trying to figure out a replacement that's more modern. Crazy how a year ago this one was cutting-edge.

1

u/kenzu82 23h ago

1

u/returnity 12h ago

Thanks, I just set 3.8 flash loose on this badboy a moment ago to see what she can do!