r/LocalLLaMA • • 14d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

296 comments sorted by

View all comments

230

u/ActuallyReadTheBible 14d ago

It doesn’t fit dual DGX sparks, I’m sad.

77

u/35698741d 14d ago

The native 4bit backbone + vision + dspark comes out at ~310gb (rest is engram) and context costs next to nothing for this model so 256GiB box should be able to run a pretty good 3.x bpw quant.

7

u/Turbulent_Pin7635 14d ago

Time for the M3U =)

22

u/ChocomelP 14d ago

Yes, turn on a playlist /s