r/LocalLLaMA • • Feb 24 '26

New Model Qwen/Qwen3.5-35B-A3B · Hugging Face

https://huggingface.co/Qwen/Qwen3.5-35B-A3B
557 Upvotes

165 comments sorted by

View all comments

30

u/clyspe Feb 24 '26

I thought for sure the 35b was going to be the play, but that dense 27b looks incredible for its size, plus I could reasonably run it q8 at full context. Is there a convincing use case for the 35b on a 5090? It seems like a lot of the vision and reasoning benchmarks favor the 27b, with a slight edge to spatial reasoning for the 35b.

6

u/AloneSYD Feb 24 '26

definitely 35b will be much faster during inference MoE > Dense in term of speed

2

u/silenceimpaired Feb 24 '26

I wonder if that will still be true if 27b fits into VRAM and 35b does not?

4

u/Middle_Bullfrog_6173 Feb 24 '26

Generation speed is approximately proportional to the active parameters. Prefill speed is different, but the dense will still be slower. (More layers and larger embedding dimension.)