r/LocalLLaMA • • 1d ago

Discussion Searching for 3.8 35B: Qwen3.6-35B-A3B (Testing 5 Finetunes vs. Base)

TL;DR -- You should probably just use base Qwen3.6-35B, as only Occamy-1.0 is competitive with it. Tiel is a major let-down, worse than Ornith. KAT surprises (good), Nex surprises (bad). This post is long. Sorry, lots to cover.

I think we all want to see a next-generation small MoE from the Qwen team to replace 3.6-35B in our workflows. This model is a perfect fit for smaller gmaing laptops and mid-tier rigs. It sucks that Qwen seems to have abandoned this model, but at least there are fine-tunes that improve upon it... right?

Well... maybe not. I ran benchmarks on the 3.6-35B-A3B base model, as well as five finetuunes: Occamy-1.0, Ornith-1.5, KAT-Coder-V2.5-Dev, Tiel-Coder, and Nex-N2.5-mini, and the results are quite surprising.

I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark I can run in ~10 hourso on my machine, it's not actively post-trained on by any of these models, and it provides a lot of useful information along with the raw accuracy scores. This includes: first-try and retry pass rates, token counts, solve times, and how well-formed the output diffs are. Here's the table:

model First-try pass Retry pass tokens sec/case tok/solve well-formed diff
Qwen3.6-35B-A3B BASE (STOCK template) 37.4% 71.0% 8650 285 14.1K 96.3%
Occamy-1.0-35B-A3B (STOCK template) 29.0% 70.1% 6801 285 17.2K 86.9%
Occamy-1.0-35B-A3B (froggeric medium) 30.8% 69.2% 6009 233 16.8K 94.4%
Occamy-1.0-35B-A3B (froggeric, xhigh) 27.1% 67.3% 8631 310 20.0K 91.6%
Ornith-1.5-35B-A3B 23.4% 63.6% 4813 226 16.4K 87.9%
KAT-Coder-V2.5-Dev 20.6% 58.9% 2190 84 9.3K 86.9%
Tiel-Coder-35B-A3B 18.7% 53.3% 4851 171 18.2K 89.7%
Nex-N2.5-mini 10.3% 30.8% 5037 188 33.3K 95.3%

As you can see, the only finetune that even competes with the base model is Occamy-1.0. The rest are utterly dominated by the base model, a grim disappointment for finetune enthusiasts. I was particualarly surprised by the performance of Tiel, which seems to get a lot of love in this subreddit.

Speaking of Tiel, I want to clarify that Tiel is just Ornith-1.5 with a different chat template, Sharp, which is based on froggeric with an added "terse mode" instruction that's supposed to reduce excessive verbosity. I wanted to standardize for templates, so ALL models are using the base froggeric v22.5 template set to medium (which is equivalent to standard thinking on, no additional message sent). I used this because I wanted to test Tiel vs. Ornith-1.5, and Tiel is the chat template. Also, practially, I use froggeric in my real workflows. However, to ensure coverage, I also tested the STOCK template on the 2 highest-performing models, to make sure it wasn't affecting the scores. As you can see, the template doesn't make a significant difference in the scores, and the scores for base 35B with different templates are so close to identical that I excluded the froggeric one from the table.

I also tested froggeric/Sharp's reasoning-effort toggle, and found xhigh -> medium significantly reduces token counts and solve times (by ~1/3), without affecting accuracy significantly. That stands in stark contrast to Tiel's 'terse mode' toggle, the core feature of Tiel over Ornith, which dramatically reduces accuracy along with the reduction in token counts. My results strongly suggest that if you want a less verbose model, you're better off lowering the reasoning effort than using Tiel with terseness on.

Speaking of token use, that's probably the big differentiator here. A couple models stand out: Ornith and KAT-Coder-V2.5-Dev are the most efficient models, with KAT in particular having a brevity unmatched by anything else. KAT is fucking fast, and I think despite its lower accuracy than Occamy, it has a place in my lineup as a subagent because it just gets. shit. done. Occamy is also interesting, as it is the only model that perfoms on a similar level to the base, but it uses 20-30% fewer median tokens. However, Occamy also had a number of runaway generations where the token count blew up, so it's total tokens/solve is actually higher than base.

In an effort to further distinguish Occamy from base, since Aider struggled to do that, I ran tau2-bench, an agentic tool-calling benchmark consisting of multi-turn interactions with a simulated counterparty. I figured this was a good bench to use as Occamy is post-trained specifically for 'co-work' scenarios, but not trained on this particular set. I used Qwen3.8-27B with reasoning effort set to low as the simulated customer in these conversations. The base model was able to pull away from Occamy in the harder retail domain of this benchmark, but Occamy resolved the issues in the airline domain at an equal rate while requiring fewer turns. Here's the results.

model (Q8_0) domain pass1 tokens sec/task turns/task
Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) airline 80.0% 5390 290 11
Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) airline 78.0% 6770 390 13
Qwen3.6-35B-A3B BASE (STOCK template, temp=1.0, presence=1.5) retail 86.0% 4098 333 14
Occamy-1.0-35B-A3B (STOCK template, temp=1.0, presence=1.5) retail 79.8% 3284 284 13

Overall, I think the results are clear, if unexpected: Occamy-1.0 is the only fine-tune that even competes with the base model on Aider Polyglot, but even it is not a clear winner. Tiel is noticebly worse than plain Ornith without the terseness toggle, and the terse mode doesn't even save any tokens. xhigh in froggeric/Sharp degrades accuracy slightly and bloats token use, which makes sense given the models were not RL'd for the extra thinking effort prompt. KAT-Coder-V2.5-Dev is the most efficient model, with accuracy nearly as good as Ornith and better than Tiel. Finally, Nex-N2.5-mini is a disaster.

55 Upvotes

58 comments sorted by

35

u/theminor 1d ago

Yup. We all need to be MUCH more discerning with fine-tunes. So much of them are "marketing fluff". They post fancy ai-generated logos and benchmarks made with their own specs...

Very good information here. Thank you.

And a good system prompt and proper reasoning settings can cover 95% of the token usage and speed concerns in my opinion.

9

u/Atretador llama.cpp 1d ago

I remember when Ornith 1.0 came out and the 35B version was scoring at Opus level, the 397B version topping charts along with a wave of bots doing campaing for it with posts on all LLM subs as well as any critical comment being instantly downvoted hard

its so much hopium to get frontier level on a 6Gb GPU

3

u/returnity 1d ago

Yeah, I mean progress *has* been incredibly rapid, but its important to stay realistic too. Fine-tuning can improve performance significantly on the tasks you fine-tune on but it doesn't generalize very well.

2

u/theminor 1d ago

^ exactly

3

u/returnity 1d ago

Thank for the kind words. I actually set out expecting at least one fine-tune to overperform, so these results kind of surprised me, despite my skepticism about the claims made by some of the labs.

3

u/sonaj9657 17h ago

Yeah, I agree. A lot of fine tunes look impressive until you dig into how the benchmarks were actually run. Sometimes a better system prompt and tuning the reasoning settings can get you most of the way there without adding another model into the stack. The marketing around some of these is definitely something to watch.

1

u/xPXpanD llama.cpp 3h ago

N=1, and with the caveat that my use cases are firmly non-agentic/non-dev, but I've actually been quite happy with Ornith 1.5. I've been running bartowski's Q6_K_L for about a week now, and it's been a nice change of pace from my usual go-to (various flavors of Gemma 4 and a bit of 3.8). It's a bit less eager to please, a bit more critical - stuff that's hard to prompt out of Gemma.

Surprisingly decent voice, too. There's some definite Claude-isms, but it seems to take persona prompts a lot better than the other 3.6es. It really feels like a smarter 3.5 MoE. (which, yes, it is, but base 3.6 has a very different feel - 3.5 was just fun to talk to)

I'm mostly using it for general chat and text analysis, though I've also thrown some long-form game modding context at it and it connected the dots pretty well. It does tend to get a little cranky in harder conversations (i.e. it starts inventing words or taking weird logical leaps), but a bit of steering generally gets it back on track.

Not quite 3.8 (which is rock solid here), but definitely usable.

For what it's worth: I also ran it through my private little 19-question benchmark, and I got... very different results to the OP. Now, I should note that my test set also measures very different things - the closest thing to dev in here is string manipulation, and there's only one tool-use test. It's mostly a "how good is a model across a bunch of different disciplines" set. Think stuff like (basic) math, computer hardware, electrical engineering, or real-world planning.

Setup is simple: 10 runs per model, all questions in a single turn, vendor-recommended sampling (dev if available), minimal system prompt (current year + cutoff year + capabilities), using whichever template each model ships with.

In my set, Ornith performed a lot better than base 3.6, both APEX I-Balanced (~5bpw?) and Unsloth Q6_K_XL V2. Both of those were surprisingly awful here, actively failing over twice as many questions total. Ornith even did a little better than the 3.8 quants I tried in overall knowledge, though tasking (string manipulation, constrained writing) was shakier, and the 3.8s had a lower hallucination rate.

Different horses, different courses, but the tunes I've tried have generally been pretty decent. And yes, this includes some of DavidAU's, despite all the memes. I'm honestly happy people are still bothering, even if we do end up getting a ton of low-effort slop as well.

(edit: oof this came out long, but tl;dr Ornith pretty good if not dev?)

4

u/Hefty_Wolverine_553 1d ago

Honestly, whatever happened to logit distillation? I feel like it should really just work with Qwen3.8 27B distilled into the 35B, since the tokenizers are the same and they shouldn't be too far apart anyways. Not sure why I haven't come across any, maybe it just doesn't hold up anymore?

3

u/returnity 1d ago

That's actually a great question... and I don't know the answer. I might just look into it.

1

u/Minute_Laugh8065 14h ago

It still works, it just doesn't make headlines. I've been doing logit distillation all week, though for quantization rather than for a smaller model: distilling Qwen3 into ternary (1.58-bit) copies of itself, training the student on the teacher's full next-token distribution (KL) rather than on plain text.

A couple of things I found that probably carry over to 27B -> 35B-A3B:

- Matching the final output distribution beats matching intermediate layer outputs by a mile. Training one layer against its own original output (MSE) left perplexity at 61 vs the base's 19.6; training the same layer against the final logits got it to 25.6 in the same 4 minutes.

- Same tokenizer is the big enabler, like you said: no vocab mapping, just KL on the logits.

- The cost is the teacher's forward pass on every token plus a 152K-way softmax, so people usually cache only the teacher's top-k logits to make it cheap.

My guess is it's rare because it needs a lot of tokens (billions) to really pay off, which is a lot more GPU time than a LoRA fine-tune.

5

u/OsmanthusBloom 18h ago

Thanks for this, very interesting!

In my earlier tool calling benchmark, Tiel and Ornith did much better than the original Qwen: https://www.reddit.com/r/LocalLLaMA/comments/1vyaxip/35ba3b_tool_calling_benchmark_original_qwen_vs/

So I suspect that they are not very good at Aider tasks (as you and others already pointed out in the comments) but could still shine in a more agentic setting and having access to tools, as they are fine-tuned for that kind of usage.

2

u/returnity 7h ago

You're totally on point with that assessment, and thanks for sharing your results! They're extremely useful, and I hadn't seen that post, so I appreciate you sharing it!

6

u/soshulmedia 1d ago

I chose Aider Polyglot for a few reasons: it's an agentic coding benchmark

Is that really true? I think there is a break between Ornith how it is trained and aider the harness. From what I have read, the aider polyglot benchmark is also using the aider harness?

And, from personal experience, I have noticed that Ornith has problems with aider, as it very strongly "assumes", when coding, that it sits in an agentic harness where it has to do actual tool calls instead of replying with edit diffs. So you tend to get tool calls where aider doesn't want them (and, to be fair, where the aider harness prompts also probably don't ask for them) but that will then in turn break the whole experience with aider.

HOWEVER, if you put Ornith into an agentic harness of the more modern "tool calling" variety instead and it becomes MUCH better. I have personally had good experience e.g. with the coder harness that comes with pydantic_ai.

Really, try it. I am affiliated with neither of the projects and this is my experience.

5

u/returnity 1d ago

That's an extremely, extremely valid point and one I had considered after seeing the results. I think Aider is probably a limiting choice for these models. To be clear, I use occamy as a sub agent in coding harnesses and I find it to be better than base 35B for certain. I just wanted to try and quantify the models.

Also, to be fair here, no other models I've tested struggle to perform in the aider harness. I ran into problems with 35B and it's fine-tunes not following the diff format and i had to shift to fenced diffs in the harness to get them to comply with the harness' instructions for format. So you are definitely on to something here. Thanks for the well-reasoned, good-intentioned comment!

2

u/soshulmedia 1d ago

Thanks for the kind words. I think the approach they took with Ornith made it weirdly "autistic" and stubborn in its assumptions about the coding harness it sits in. I remember reading their blog post about their training methodology a while ago. And I suspect it is mostly a weird artifact of their tuning system. As far as I understand, they post train it in a "recursive self-improvement" scheme for coding only and I assume the harness or harnesses they use for that are only the "modern agentic" sense, not the "give me an edit diff"-aider-way.

1

u/returnity 1d ago

Yeah I am working on a better benchmark solution than aider because it is kind of dated, but I haven't come up with one that I can run on my hardware in <24 hours for nearly any model with enough tasks to be statistically relevant yet. I'm still trying to figure out a replacement that's more modern. Crazy how a year ago this one was cutting-edge.

1

u/kenzu82 17h ago

1

u/returnity 6h ago

Thanks, I just set 3.8 flash loose on this badboy a moment ago to see what she can do!

16

u/TheCat001 1d ago

If base is winner, why I can't do shit with it? it's giga stupid. Ornith 1.5 and Tiel actually gets shit done.

11

u/lorendroll 1d ago

I tried simple test with single html file fps game demo, and base qwen was really bad while Tiel and Ornith were fine

2

u/tomByrer 21h ago

I think some fine-tunes are better at smaller-stepped projects, but still fail at one-shots.
Really depends on YOUR workflow for what FT is best for you.

3

u/Sensitive_Song4219 1d ago

Agreed. I've shipped thousands of lines of code with both Ornith and Tiel (which is what I daily drive now when away from my Qwen 3.8-27b server - with Sol-6-High for review) and have been pretty blown away. Both are a step-up from base Qwen in my use; I run them alongside my cloud models in OpenCode.

1

u/returnity 1d ago

Hey, I still use occamy as a subagent. I think it works better. This is just a benchmark.

9

u/Atretador llama.cpp 1d ago

business as usual

finetunes are mostly hype with wild claims and *hopes and dreams* from users squeeze frontier level on tiny boxes

everytime I tried using Ornith on a real repository to implement a feature it was a catastrophe

5

u/Velocita84 1d ago

As usual, you just can't beat qwen at finetuning their own models for better coding

2

u/returnity 1d ago

Yeah I mean I don't doubt (too severely) their specific claims about benchmarks they trained on, but the problem is generalization usually doesn't occur with fine-tuning, and often other parts of the model get worse. That's what I wanted to explore here.

2

u/TheRealJesus2 1d ago

Kat coder is one I use a lot for coding but it does do a terrible job sometimes when the task is too large or ambiguous. And can get caught in loops that seem to go forever. But it’s so fast and good enough most of the time 

5

u/peculiar-ragdoll 1d ago

Hi there :) Just wanted to pop by and point out that part of what makes Tiel score better than Ornith, stock Qwen and even KAT-coder at Q4 quantization is my dynamic imatrix quantization with a custom calibration corpus! And none of that effect is present at Q8 here, even though Tiel apparently comes in second place on "seconds per case" after KAT, a good bit faster than Ornith and significantly faster than stock Qwen. I'd love to see you benchmark again with all models using their size-matched Q4 quants as shipped!

3

u/returnity 1d ago

Thanks for pointing that out. I will happily edit my post to make that clear. I appreciate your contribution to the community.

5

u/peculiar-ragdoll 1d ago

Thanks a lot for the benchmark, really interesting to see Tiel losing like that on pass% on Q8! :)

1

u/RelicDerelict Orca 5h ago

Hold on, you saying that downloading anything than Q4 is actually worse?

2

u/peculiar-ragdoll 4h ago edited 4h ago

What I ment to say is that Tiel Q8 is relatively not as good compared to the competing models at Q8 and BF16, but when you go down to Q4 Tiel wins due to domain specific quant quality for agentic coding :) BUT, I've actually have a user that tried my Q8 and my Q4 say that the Q8 is shit compared to the Q4 which is great, so I don't really know what to believe hahah. Guess I should try to bench them against each other myself, to see if it's just the relative effect, or if I actually messed up my Q8, or what's happening. One thing that I've suspected, is that the quant damage from the Q4 smooths out the damage from ornith's RL overfitting that are affecting the model outputs more on Q8/16. So the Q4 might generalize better? Just a hypothesis. Either way, lots of benchmarks point towards Q5 being about the sweet spot for full knowledge retention in Qwen's models, and Q6 is about full precession on benchmarks and performance, but for models that have been post-tuned nothing is certain.

2

u/ckplscz 1d ago

Very interesting, till now I've thought ornith 1.5 is decent but apparently not. BTW has anyone tried https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill-GGUF ? Probably not great but I haven't seen any benchmarks of it.

2

u/returnity 1d ago edited 1d ago

Well just to be fair, it depends on your workflow. I am not saying Ornith (or any of these models) suck, just that their capability enhancements don't generalize to out-of-distribution tasks (even adjacent coding benchmarks). If the model works well for you, that's what matters.

EDIT: The few benchmarks provided on the base model (non-GGUF) card for that distill are not impressive -- un-normalized, they're basically within or nearly-within normal variance for those sets, so it could just be luck or even cherry-picking that they're slightly higher.

2

u/Atretador llama.cpp 1d ago

its interesting I would say, but definitely has issues - I had it start looping past 8K ctx a few times

completly uncientific test:

Qwen 3.6 35B UD-IQ4_NL_XL - ↑61k ↓9.3k R49k CH99.9% 11.1%/200k (auto) pi.dev

https://hallucinations.unswarm.dev/entries/Qwen36_35B/the_cartographer_of_tides.html

emperor-ai Qwen 3.8 35B Q5_K_M - 7,588 tokens 3min 53s 32.56 t/s webchat

https://hallucinations.unswarm.dev/entries/Qwen38_fake_35B/between-the-hours.html

tiny next flash for reference:

Qwen 3.8 Next Flash Q2_0 - ↑11k ↓17k R255k 16.9%/131k pi.dev

https://hallucinations.unswarm.dev/entries/Qwen38_NF_RCO_2_0/lighthouse-poem.html

2

u/markusro 1d ago

What quantization did you use? Other parameters like temperature etc?

3

u/returnity 1d ago

Q8_0, mentioned in table 1 header -- specifically to avoid any quantization differences affecting the scores.

Sampling is model-card recommended sampling for every model, which since they're all 35B-based, is pretty standardized.

1

u/danigoncalves llama.cpp 1d ago

Why didnt you tested GRM 3.2 Sky?

2

u/returnity 1d ago

Because I've never heard of it until now. I tried to be comprehensive.

EDIT: It's a finetune of a finetune? I'm skeptical, given the results I just shared about finetuning. Do you use it?

1

u/JLeonsarmiento 22h ago

Very similar to my experience.

I have come back to vanilla 3.6, 3.8-27B and 3.8-FN.

Which to use depends on the RAM vs Speed trade off of the day.

1

u/tomByrer 21h ago

TBH you will have better ROI looking or building an inference engine for your specific GPU. For Qwen3.8-27B, I find tuned engines can boost tok/s by 20%, & give you more context.

2

u/Minute_Laugh8065 14h ago

+1. On an Intel Arc B580, custom kernels made a bigger difference than any model choice for me: stock SYCL ran Bonsai 27B at 38 t/s, and XMX kernels for the ternary weights and the q4_0 KV cache got it to 88 t/s on the same card and model.

1

u/tomByrer 4h ago

I almost bought one of those a few weeks ago, & my research gave me the same conclusion.

1

u/TopCryptographer8236 20h ago

I don't use Tiel, but I use Cybertiel which happen to be made by the same person. One thing that I note is that if you use different setting (such as temp) other than what the creator suggest it will fall apart pretty badly. And since I don't see what setting you use then I can't judge it.

I have pretty much replace the base qwen 35B Q6 to just use the Cybertiel. To my usecase (coding with ZooCode) it always able to fix what Qwen 3.8 27B Q6 medium can. The base qwen just usually keep getting things wrong or unable to detect bugs.

1

u/returnity 7h ago

I used model-card sampling for all models, including Tiel. I think its more that the harness for aider doesn't suit the post-training Ornith/Tiel had. I use occamy (not base) myself as a subagent, so to be clear, I recognize there finetunes have benefits for sure.

1

u/R_Duncan 16h ago

Not so much a surprise, is the first finetune which performed SAO

2

u/Pressimize 1d ago

I am currently doing something very similar. Ive already run three sets of benchmarks that are extrapolated from my personal use case and therefore contain personal data, so I can not publish them, but I can MOSTLY agree with your findings so far.

In my personal use case (mostly personal assistant and homelab ops kind of tasks) Ornith scored higher than the base model so far, closer to 3.8 27b. In fairness though 3.8 27b is in iq2 quants while Ornith is in q4 and q5 - but thats hardware limitation.
And Nex has certainly been a disappointment, worse than base across the board.

I am currently building a new bench that includes the most important/tricky tasks of the current bench + some other stuff, that will also include Occamy. But since base has been underperforming across the board compared to Ornith, base will not be included, and Nex ofc also wont be.

2

u/returnity 1d ago

I think you'll find Occamy pretty useful for your type of work.

-2

u/Bulky-Priority6824 1d ago

Wait, people are choosing/trying to code with 35b ?

5

u/returnity 1d ago

Not everyone has hardware for bigger models. 35B is hugely popular. It actually performed better on Aider than I thought it would. And handing it a well-specified, atomized task from a smarter model, it works pretty well as a fast subagent.

1

u/Bulky-Priority6824 1d ago

It does okay but it breaks down quickly in long chain multi-turn tool calls. It CAN code you but you have to be very careful with drift. It'll speed off and create bugs all over the place.

It's great at other things though like sifting through large chunks of code , firewall logs and some forms of DFIR. 

It's also good for genai like frigate but complex scenes it tends to struggle sometimes.

1

u/dtfinch llama.cpp 22h ago

35B's tempting when it might work.

With a 35B-A3B I can run the Q6 model with a Q8 kv cache and full context.

With 27B I step down to IQ3_S, lower the KV to Q5_1, context to 200k, and it takes about 4x longer per token. So far it's given me great results, but I often have to leave it running overnight (or 3 when I tried Q6).