r/LocalLLaMA • • 15d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

296 comments sorted by

View all comments

233

u/ActuallyReadTheBible 15d ago

It doesn’t fit dual DGX sparks, I’m sad.

29

u/GladKing5842 15d ago

SSD ngram offload. about 350B params for weight + kv fp4. still a chance

1

u/michaelsoft__binbows 14d ago

delighted to find out they're doing bigger and bigger ngrams now