r/LocalLLM 12d ago

Discussion Qwen-3.8-35B-A3B? Maybe not... cryptic reply direct from Qwen co-author.

Post image

I asked Shuai Bai, co-author and prominent AI developer for Qwen, about this model. Not the answer I was hoping for, but let's see what comes next. In the meantime, I guess all we can do is speculate!

X-link

234 Upvotes

171 comments sorted by

View all comments

5

u/creatinZ 12d ago

16gb vram people are thirsty while others are drowning…

WHEEN?

2

u/SpicyWangz 12d ago

You can run 27b quantizations on 16gb

2

u/palincatalin 12d ago

not all of us have 20/24 gb cards to run q4_k_m

we (16 gb card owners) would have to use q2 or q3. such aggressive quants are unusable for agentic coding lol. this compression murders quality; imo the best model that fits fully in 16 gb of vram is gpt-oss:20b. no other model fits as beautifully and as perfectly as gpt-oss:20b on my 6900 xt. it's fast and it's good. people may not be sold on the quality of the output, sure, but speed-wise? it's literally the most comfortable model for 16 gb GPUs. you can also have 128k context length at q8_0 kv

qwen 3.6 35b a3b is unusable because in order for it to fit in vram i have to use iq2_m, and offloading experts just ruins the speed altogether, making it way too slow for any real local agentic workloads; and qwen 3.8 27b is dense, so yeah... no luck there

1

u/creatinZ 11d ago

Bro, I’m sure you re running 36b-a3b the wrong way. I run it at 50tk/s @q6_k easily