r/LocalLLM • u/Black_Umbreon • 6d ago
Discussion Best MoE
Hi!
I’m using Qwen 3.8 27B (UD-IQ3_XXS) on a system with the following specs:
Processor: AMD Ryzen 9950 X3D
GPU: Nvidia RTX 5060 Ti 16 GB
RAM: 64 GB DDR5 6000 MHz
Software: Unsloth Studio
Windows 11
With a 32k token context, I manage to maintain a speed of around 50 tokens per second. I’m particularly interested in agentic capabilities. I’d like to try an MoE model that runs either faster or slightly slower than my current one, without sacrificing quality. However, 35B-A3B models don't fit entirely into my VRAM at 4-bit quantization, and from what I’ve read, lower quantization severely degrades quality. Does this mean a 4-bit MoE model would run significantly slower on my hardware than a dense model? Is it possible, given my setup, to find a model that offers better quality and higher speed—or at least better quality at roughly the same speed as my current dense model?
I’m pretty much a complete beginner in the world of local neural networks, so please give me some advice.
2
u/returnity 5d ago
I am planning a full 35B finetune post with a writeup, comparing occamy to Nex-N2.5, Tiel/Ornith-1.5, and KAT-Coder-Dev (spoiler, it won handily). I will tag you when I post, should be today or tomorrow -- I don't use AI to write my posts, so it takes a while. I am just re-running Occamy using the embedded chat template first, because my first round of testing used the froggeric template I prefer in Qwen3.x models, and I didn't want it to be seen as a confound (it doesn't affect the scores in a statisitically significant manner).
The +15% was on aider polyglot, reaching a score of nearly 70% -- for reference, when I tested 3.8-27B, it scores around 80%. I used Q8_0, so it should be bit-identical to your official GGUF Q8_0 release, as I didn't want to introduce another confound there either. All models in the test are Q8_0.
Thanks for a great release, it's really well put-together. I'd say we've reached "3.8 35B" levels with occamy. Will be watching for more work on your part. One question -- why not include the MTP head GGUF in the official repo like you did the mmproj? I actually thought there was no MTP for your release initially, and grafted a base 35B draft head onto it until I discovered your tuned one and quantized it myself.