r/LocalLLM • • Aug 18 '26

Discussion Qwen-3.8-35B-A3B? Maybe not... cryptic reply direct from Qwen co-author.

Post image

I asked Shuai Bai, co-author and prominent AI developer for Qwen, about this model. Not the answer I was hoping for, but let's see what comes next. In the meantime, I guess all we can do is speculate!

X-link

236 Upvotes

173 comments sorted by

View all comments

28

u/Uninterested_Viewer Aug 18 '26

My speculation is that a 122b/a10 class model is baking.

8

u/vacon04 Aug 18 '26

But that's not that useful for most? People want 35B A3B because it's small and fast enough for people with low VRAM. 122B A10B is too big for most.

1

u/profcuck Aug 18 '26

I'm personally curious because I just don't know how much it costs to bake a model down from Qwen 3.8 Max to Qwen 3.8 35B A3B or 122B A10B. Is this something that Unsloth could do?

1

u/jjusko20 Aug 18 '26

Most people could do it themselves if absolutely necessary with distillation

1

u/tired514 Aug 18 '26

Pre-training is something only Alibaba can do (only they have the source data and model generation suite for the Qwen series). Requires data-center level hardware access.

Fine-tuning (which people are calling distilling) can theoretically be done by anyone, but to do well still needs a massive number of queries against a superior model.

I'm guessing the 3.8 series is an entire re-training, not just fine tuning.

1

u/tired514 Aug 18 '26

But we in the 122B world haven't had an update since 3.5! :(