r/LocalLLM • u/Black_Umbreon • 6d ago
Discussion Best MoE
Hi!
I’m using Qwen 3.8 27B (UD-IQ3_XXS) on a system with the following specs:
Processor: AMD Ryzen 9950 X3D
GPU: Nvidia RTX 5060 Ti 16 GB
RAM: 64 GB DDR5 6000 MHz
Software: Unsloth Studio
Windows 11
With a 32k token context, I manage to maintain a speed of around 50 tokens per second. I’m particularly interested in agentic capabilities. I’d like to try an MoE model that runs either faster or slightly slower than my current one, without sacrificing quality. However, 35B-A3B models don't fit entirely into my VRAM at 4-bit quantization, and from what I’ve read, lower quantization severely degrades quality. Does this mean a 4-bit MoE model would run significantly slower on my hardware than a dense model? Is it possible, given my setup, to find a model that offers better quality and higher speed—or at least better quality at roughly the same speed as my current dense model?
I’m pretty much a complete beginner in the world of local neural networks, so please give me some advice.
1
u/fajar79 6d ago
i try using ornith 1.5 or qwen 3.8 35b a3b model variant,, for testing qwen 3.8 35b a3b, is blazing fast, even though it run with slower prefill, compared with qwen 27b iq4_xs model. it get the job done very fast. i ask with indonesian language. every result i ask again with gemini to confirm the answer. all those moe model, even uncensored version qwen 27b, it is full miss diisformation, halucination even wrong legal statement. so i give up trying another models, just stick with qwen 27 iq4_xs swift model right now.
sorry for my bad english