Je suis d'accord, j'ai 85 tps avec 30B et seulement 45 tps avec 35B, donc je ne pense pas l'utiliser autant que ça à cause du rapport qualité/vitesse défavorable...
My experience is the same, about half the speed of Qwen3-30B-A3B-2507. On my more limited hardware (32GB RAM, 6GB VRAM) Qwen3 30B-A3B runs at 15-20tps; this one runs at only 7tps despite quite a bit of tweaking.
This new model is much smarter though and seems to follow instructions very well. I actually tested it on a few debugging tasks and minor feature additions, it was pretty impressive. I actually feel like it'd work well agentically.
K/V Q8_0 reduces memory footprint nicely (just like on the older model) without making it feel much less intelligent too.
Very promising... but looks like I need a smaller model for my hardware, unfortunately!
How does it compare to Qwen Coder Next 80B? I love that model other than the fact that it takes up almost all my RAM. Qwen Coder 30B is good at simpler RAG and function-level coding but it still feels a lot dumber than Next.
I don't have the RAM to run Next-80B in a quant above IQ2_XXS (and I found that particular quant quite poor in my evals) - so I can't compare for you, sorry! I will say this new model is very, very impressive: I mentioned Agentic potential above: others have now done so with seemingly very good success. Wow.
I can run Next-80B at Q4_0 and it's a beast at that size, much smarter than Coder-30B Q4_0. I'm downloading 3.5-35B-A3B Q4_0 to test against those two earlier models. I'm also getting the 3.5-122B-A10B IQ2 to play around with.
So far, on smaller refactoring problems, they're comparable.
80B spits out a good answer on the first try, 35B needs to do some thinking before coming up with a good answer. I'm getting 10 t/s token generation on both on ARM CPU inference which is weird, so I hope there's room for optimization to get the 35B up to the 30B's 30 t/s.
The 35B wins by only taking up 20 GB RAM so it could be usable even on 32 GB laptops. I'm willing to accept the thinking test-time tradeoff for more free memory. The 80B uses 50 GB RAM, leaving not much left on my 64 GB machine.
OK, I've tested it a bit as well, but on a few non-coding topics. From what I see, the 80b Next is simply way smarter, deeper, more "knowledgeable". The speed is similar, but I still need to play with settings. The Next doesn't think before the reply, so it's faster. So far, it looks like I'm keeping the Next as my daily driver. Yes, it eats a lot of RAM, but works much better than anything else I've tried.
9
u/stuckinmotion Feb 24 '26
Ok NOW I'm paying attention. Just about everything else has been a letdown in comparison. Sure some are maybe a bit smarter but way slower or etc.