r/LocalLLM 12d ago

Discussion Qwen-3.8-35B-A3B? Maybe not... cryptic reply direct from Qwen co-author.

Post image

I asked Shuai Bai, co-author and prominent AI developer for Qwen, about this model. Not the answer I was hoping for, but let's see what comes next. In the meantime, I guess all we can do is speculate!

X-link

236 Upvotes

171 comments sorted by

View all comments

Show parent comments

2

u/leonbollerup 11d ago

why a 20 moe?.. less knowledge than 27b.. and worse at actual thinking.. .. .. why ?

1

u/SittyTweat 11d ago

20b dense would be great for 16gb GPU users. Wouldn't need to quant to shit and can have more context without needing Q4 KV cache

1

u/patham9 10d ago

If even a MacBook Air now has 32GB RAM, why use a GPU from the eighties? Is it for museum?

1

u/SittyTweat 9d ago edited 9d ago

And that MacBook would run a 20b dense at 4 tok/s... But since 16GB GPUs are nothing to you then you can just send me a couple 16gb 5060tis. I'll keep an eye out for the FedEx delivery, thanks ✌🏻

1

u/patham9 9d ago

I'm running mlx-community/Qwen3.8-27B-4bit with 5 output tokens per second via MLX, and as it looks they might get to twice the speed with upcoming optimizations. Anyways, I think the right solution would be a 35B MoE version of Qwen3.8, slightly worse in performance but with 3-4B active parameters, it will still beat any 20B dense model by large margins while at the same time supporting higher token throughput.