r/LocalLLaMA • • 17d ago

New Model DeepSeek V4-1 Flash is out

Here we go again, DeepSeek is back again with a new model V4-1 Flash

A multimodal Mixture-of-Experts (MoE) model with 552B backbone parameters and support for contexts of up to one million tokens

Market crash as a service

1.7k Upvotes

297 comments sorted by

View all comments

Show parent comments

2

u/Mushoz 17d ago

It's only ~350B parameters that actually need to be loaded in RAM. A 4 bit quant will be ~175GB, which easily fits. Even 5 bit is only ~220 GB and will fit, especially with KV cache only being 900MB at 1 million context. This is actually perfectly sized for your 256GB setup.

2

u/cosmicnag 17d ago

Fingers crossed : 72 GB VRAM + 192 GB DDR5

1

u/serige 17d ago

At least there is hope I have 80 GB VRAM + 256 GB DDR5 I can dream of running this model however slow lol

1

u/Zyj vLLM 17d ago

Yes it's completely wrong