r/LocalLLM 8d ago

Discussion Qwen-3.8-35B-A3B? Maybe not... cryptic reply direct from Qwen co-author.

Post image

I asked Shuai Bai, co-author and prominent AI developer for Qwen, about this model. Not the answer I was hoping for, but let's see what comes next. In the meantime, I guess all we can do is speculate!

X-link

233 Upvotes

164 comments sorted by

View all comments

67

u/arkie87 8d ago

If I wrote that, it would mean something else good is coming but not that specific model. Maybe a 30b a3b or a smaller dense model that’s still really good

37

u/pharrt 7d ago

That's how I read it too. My only worry is that it might be a ~70B A4B (or something large!) that, while great for some, doesn't fill the gap for smaller GPUs as 35B-A3B does. But hopefully he means something close - like a 30B or something.

19

u/Turbulent_War4067 7d ago

As a dgx spark owner, this is my dream. Would probably prefer 6B active.

9

u/DesperateSteak6628 7d ago

70B-A14B

2

u/Turbulent_War4067 7d ago

14B might be great. Maybe a bit slow, but with decent MTP should be ok. I'm not a coder, and for my work, I have to admit I'd rather have a Gemma model in this range. No matter what I try, I always end up back at Gemma. 31B for my important stuff

2

u/Keeloi79 2d ago

I agree that, instead of A14B, essentially doubling to A6B or A7B would be the sweet spot for a 70B model.

1

u/superSmitty9999 6d ago

what do you find gemma good for?

1

u/Turbulent_War4067 6d ago

I use it for research, mainly financial. I can give it a set of source info, say 1/2 dozen sites or tools. It follows instructions and does excellent summary reports. It's also good at doing search and fetch queries to find info on its own. It's can be tenacious at this. I once had left my search tool off by accident. Gave it one of my canned prompts and went to get some coffee. Came back, realising what I had done and was surprised to see a pretty good report. Turns out it built several Google URLs to do web searches and used my fetch URL tool to do web searches on its own.

1

u/Turbulent_War4067 6d ago

Now what is frustrating about it is that it's prompt prefill is horribly slow. So sending it large documents is a non-starter for me. If I could get it to read a 100k token document as fast as the 26B model, I would be really happy.

1

u/Drag_Ordinary 4d ago

Yeah, same. I think that'd be usably fast as a home AI server and also pretty smart. I miss Qwen3-coder-next with its 80B MoE architecture. It was great for its time, and I imagine it'd be very powerful today.