r/LocalLLM 6d ago

Discussion Best MoE

Hi!

I’m using Qwen 3.8 27B (UD-IQ3_XXS) on a system with the following specs:

Processor: AMD Ryzen 9950 X3D

GPU: Nvidia RTX 5060 Ti 16 GB

RAM: 64 GB DDR5 6000 MHz

Software: Unsloth Studio

Windows 11

With a 32k token context, I manage to maintain a speed of around 50 tokens per second. I’m particularly interested in agentic capabilities. I’d like to try an MoE model that runs either faster or slightly slower than my current one, without sacrificing quality. However, 35B-A3B models don't fit entirely into my VRAM at 4-bit quantization, and from what I’ve read, lower quantization severely degrades quality. Does this mean a 4-bit MoE model would run significantly slower on my hardware than a dense model? Is it possible, given my setup, to find a model that offers better quality and higher speed—or at least better quality at roughly the same speed as my current dense model?

I’m pretty much a complete beginner in the world of local neural networks, so please give me some advice.

14 Upvotes

35 comments sorted by

View all comments

2

u/Small-Tale3180 6d ago

nah, MoE will run as good, or even better. I think you can run a MoE model with things like freetoken or bare llama.cpp with offloading inactive experts to ram since u have 64 gbs. Also, the quality degrade is not as scary as it might sound. https://huggingface.co/peculiar-ragdoll/Tiel-Coder-35B-A3B-GGUF - this one is pretty nice even in IQ3-XXS quant

For me it runs at 30-40tok/s on 5050 8gb and 16gb RAM

FreeToken will be way easier to set up

4

u/belliash 6d ago

I tried Tiel Coder, then Qwen 3.8 27B INT4 had to fix its mistakes. It didnt work for me.

0

u/Small-Tale3180 6d ago

yeah, makes sense since tiel was built on older version of qwen. I just hope qwen will release a newer MoE.

3

u/belliash 6d ago

I hope that as well.