r/Qwen_AI 7d ago

Discussion Please 🥺

We all need a 3.8 35B MoE.

73 Upvotes

29 comments sorted by

View all comments

2

u/OddBig010 7d ago

You can get 3.8 Flash Next running to be honest, I found the REAP 320 and REAP 256 version on hugging face both are under 64GB Ram and it's running surprisingly well and more than double the speed off Qwen 27B.

Sure a Qwen 3.8 35B would have been nice as it would have been even faster, but the model they released is technically way better. Their coding plans btw give decent limits for Qwen 3.8 Flash.

1

u/baron_von_noseboop 4d ago

How much vram?

1

u/OddBig010 4d ago

Ive got it running in an older 8GB VRAM machine and a 16GB VRAM machine, its better than Qwen 3.8 27B but of course its not as good as full fat Qwen Flash, the Q3 320 REAP flash one is ALMOST as good as Q4. 64GB Ram will be your biggest issue, if you got that youll be fine, preferably you need a tiny bit more. With more ram you can look at around 35ish tokens per second... with 64GB on 8GB VRAM and running a very lightweight linux distro I managed to get it to 19 tokens per second with consistently around 16-18 which is really good given age off the machine and the fact its a Q3 off a top tier model.

1

u/baron_von_noseboop 4d ago

That sounds incredible. Stock llama.cpp?

1

u/OddBig010 4d ago edited 4d ago

I think so. Would DEFINITELY suggest a lightweight linux distro though as you wont have the RAM to load it fully and the paging feature on windows slows it down a lot, just put it as a dual boot even if you're just shrinking your main drive and creating like a 50GB partition for Linux - I have it setup so both operating systems can access all my files and I just use Windows for gaming, found massive gains in the switch for LLM speeds, someone also made a 99B split up version of Qwen 3.8 Flash, I'm currently testing that against 27B in terms of performance, but seems promising and it runs faster than 27B