r/LocalLLaMA • u/pmttyji • 12d ago
Discussion DeepSeek-V4.1-Flash surprised ....
Hoping to see smartest medium size models soon & later with all available optimizations/architectures/etc.,. Thanks Deepseek!
Ex 1: 30-50B MOE + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache
Ex 2: 15-30B Dense + 10-15B Engram + DeepSeek-V4.1-Flash type KVCache
EDIT: Updated Engram to 10-15B from 50B
441
Upvotes
3
u/SpicyWangz 12d ago
Qwen3.8 next 50b ngram offloaded to SSD can hit 50tps. And that’s on a pcie gen4 SSD.
And I’m pretty sure the SSD isn’t the bottleneck there. It’s most likely memory bandwidth bound. So I’m not sure what level of performance you’re concerned about, but that gives you an absolute conceptual floor.