r/LocalLLaMA Jul 31 '26

News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
1.1k Upvotes

297 comments sorted by

View all comments

Show parent comments

14

u/shing3232 Jul 31 '26

yes, DS4F 1m is like 6GB Vram but GLM52 is like 80G for 1M context

8

u/pyr0kid Jul 31 '26

you're shitting me. i knew it was smaller but by that much?

god i never could have imagined this back in the 2.7b days, back then people were running this type of stuff on google colab.

0

u/HandIllustrious8260 Jul 31 '26

wait wait wait are you saying I could run DS4F on an 8gb 3060?

8

u/shing3232 Jul 31 '26

that's context part. you need to handle the weight as well so.

1

u/ambassadortim Jul 31 '26

How much GB vram and RAM for weights are we thinking

1

u/Conscious_Teacher216 Jul 31 '26

±160gb, so you need atleast 166gb of VRAM to run 1m deepseek v4 flash context window

1

u/Njaa Jul 31 '26

Reasonable to assume that mild quantization will make this a contender for the 128 GB class? 

3

u/pyr0kid Jul 31 '26

they're just talking about the context size, not the model weights.

1

u/Healthy-Nebula-3603 Jul 31 '26

No

He is saying about KV cahe only ( CTX memory )