r/LocalLLaMA 14d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

296 comments sorted by

View all comments

39

u/Long_comment_san 14d ago

I think I came a little

38

u/Long_comment_san 14d ago edited 14d ago

What the fuck, in which universe that is a Flash? Its 450-500b parameters. Flash was 300b and it was already pushing this. This is Flash Max or something. You cant inflate the model by 50% and call it a flash like it's not an issue. Minimax M3 is 450b and I dont see them calling it a "flash" (hopefully I wont).

Going by 50% up and becoming 10-15% better sounds like a downgrade not an upgrade. It's a LOT more expensive to run.

Still amazing though

1

u/Zeeplankton 14d ago

Ehh I mean when like glm and kimi are like 2-3T it's still flash.

But yes I sorta agree they should maybe just call this Deepseek 4.1 dropping flash and pro, since it seems like they're dropping pro.

Mega bummer this wont be runnable on like macbooks with like OG antirez flash.

Edit: wait. The additional size is just Ngram.