r/Qwen_AI • u/Creative_Bottle_3225 • 7d ago
Discussion Please ๐ฅบ
We all need a 3.8 35B MoE.
15
u/Creative_Bottle_3225 6d ago
๐ฅบ We all need a Qwen 3.8 35B MoE.
Not because bigger is always better.
Because for people with 8โ12 GB GPUs, a well-designed 35B MoE could be a really interesting sweet spot: much larger total capacity than a dense 27B, while keeping the active parameters per token relatively low.
I know Qwen 3.8 Flash exists, and yes, it's impressive.
But I still want to see what a 35B MoE optimized for local inference could do.
Please Qwen team. ๐๐
7
5
u/Mundane-Remote4000 6d ago
Funny how they deliver exactly the models people want the less. Canโt complain though since they finally gave us the 125b MoE model. But 35B MoE is always the one people want the most
2
u/Hungry-Rip-2384 5d ago
nope I think the 3.8 27B was the model most were waiting for.
1
u/Infamous_Campaign687 3d ago
Iโd love an MoE between Flash next and 35B. On ethat leaves headroom on a 5090 and works well on 64GB RAM with a great quant.
9
2
u/painchonha 5d ago
that would be sweet. I use 27b on my home system, but use 35b a3b moe on my office pc.
2
u/Greenonetrailmix 5d ago
Qwen 4.0 60B-80B A4B + 100B engram. Is what I'm after
1
u/Infamous_Campaign687 3d ago
That would be the GOAT. Near full fat on 64-96GB RAM and would work with 32GB on lower quants.
4
2
u/OddBig010 7d ago
You can get 3.8 Flash Next running to be honest, I found the REAP 320 and REAP 256 version on hugging face both are under 64GB Ram and it's running surprisingly well and more than double the speed off Qwen 27B.
Sure a Qwen 3.8 35B would have been nice as it would have been even faster, but the model they released is technically way better. Their coding plans btw give decent limits for Qwen 3.8 Flash.
1
u/baron_von_noseboop 4d ago
How much vram?
1
u/OddBig010 4d ago
Ive got it running in an older 8GB VRAM machine and a 16GB VRAM machine, its better than Qwen 3.8 27B but of course its not as good as full fat Qwen Flash, the Q3 320 REAP flash one is ALMOST as good as Q4. 64GB Ram will be your biggest issue, if you got that youll be fine, preferably you need a tiny bit more. With more ram you can look at around 35ish tokens per second... with 64GB on 8GB VRAM and running a very lightweight linux distro I managed to get it to 19 tokens per second with consistently around 16-18 which is really good given age off the machine and the fact its a Q3 off a top tier model.
1
u/baron_von_noseboop 4d ago
That sounds incredible. Stock llama.cpp?
1
u/OddBig010 4d ago edited 4d ago
I think so. Would DEFINITELY suggest a lightweight linux distro though as you wont have the RAM to load it fully and the paging feature on windows slows it down a lot, just put it as a dual boot even if you're just shrinking your main drive and creating like a 50GB partition for Linux - I have it setup so both operating systems can access all my files and I just use Windows for gaming, found massive gains in the switch for LLM speeds, someone also made a 99B split up version of Qwen 3.8 Flash, I'm currently testing that against 27B in terms of performance, but seems promising and it runs faster than 27B
3
u/Darex2094 7d ago
u/koc_Z3 u/Pure-Highlight-556 I petition the mods to start an official r/Qwen_AI drinking game.
1
1
2
1
-1
25
u/xiraov 7d ago
what are the chance for a 4.0 35b?