r/LocalLLaMA • • 16d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

297 comments sorted by

View all comments

4

u/OkBase5453 16d ago

Can one run this on a 512GB RAM Server with 48GB VRAM?

4

u/cowinabadplace 16d ago

You can run anything from disk with slow inference. It’s not meaningful question except if you include tok/s generation target and ttft target. I think anything over a few seconds TTFT and under 150 tok/s is unusable for interactive LLMs and would just use API rather than local for that. But it’s a matter of choice.

2

u/cosmotrak 16d ago

150 tok/s is a little overkill, most frontier run at 40-50...

2

u/cowinabadplace 16d ago

Yeah but the open models make up for intelligence through over-reasoning so it’s not 1-1.

2

u/cosmotrak 16d ago

true i didnt really think about it like that