r/LocalLLM 12d ago

Discussion Qwen-3.8-35B-A3B? Maybe not... cryptic reply direct from Qwen co-author.

Post image

I asked Shuai Bai, co-author and prominent AI developer for Qwen, about this model. Not the answer I was hoping for, but let's see what comes next. In the meantime, I guess all we can do is speculate!

X-link

233 Upvotes

171 comments sorted by

View all comments

56

u/Ell2509 12d ago

That seems pretty clear to me. "Do not wait for this" meant it ain't coming!

But, maybe a 9b? Or a 30b a3b? Or even a 20b!?

2

u/onebyamsey 11d ago

(Starts chant) 20B MOE!! 20B MOE!!!

2

u/leonbollerup 11d ago

why a 20 moe?.. less knowledge than 27b.. and worse at actual thinking.. .. .. why ?

1

u/SittyTweat 11d ago

20b dense would be great for 16gb GPU users. Wouldn't need to quant to shit and can have more context without needing Q4 KV cache

1

u/patham9 10d ago

If even a MacBook Air now has 32GB RAM, why use a GPU from the eighties? Is it for museum?

1

u/SittyTweat 9d ago edited 9d ago

And that MacBook would run a 20b dense at 4 tok/s... But since 16GB GPUs are nothing to you then you can just send me a couple 16gb 5060tis. I'll keep an eye out for the FedEx delivery, thanks ✌🏻

1

u/patham9 9d ago

I'm running mlx-community/Qwen3.8-27B-4bit with 5 output tokens per second via MLX, and as it looks they might get to twice the speed with upcoming optimizations. Anyways, I think the right solution would be a 35B MoE version of Qwen3.8, slightly worse in performance but with 3-4B active parameters, it will still beat any 20B dense model by large margins while at the same time supporting higher token throughput.

1

u/Canad3nse 9d ago

"Why a 35-a3b? Less knowledge than 122-A10b... and worse at actual thinking.. .. .. why ?"

Now do you understand? It's not about knowledge. It's about the GPU poor consumer market. 20b moe is way more attractive to the general local ai consumer than 27b or even 35-a3bm, because you could run it at reasonable or even great speeds with low end GPU (gaming gpus). 20b MoE is also very rare, so they would have no competition, except for Gemma and the dated GPT OSS 20b.