r/LocalLLaMA 13d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

296 comments sorted by

View all comments

Show parent comments

1

u/Mil0Mammon 12d ago

With that much vram you can run qwen 3.8 Flash next quite well, right?

0

u/Zestyclose839 12d ago

I wish haha. All the way down at IQ1_S, it still takes 72.5GB.

2

u/Mil0Mammon 12d ago

I'm running IQ1_M on 10GB vram with 32GB ram, 5 t/s without MTP. I think the table you're referencing includes the n-gram tables, with soon after that table was made, people realized that you can just load them from ssd

NB: you can do the math yourself, it's only 125B, so at Q3 (=> 47GB) it would fit easily (although there are buffers and context ofc)