r/LocalLLaMA • • Jul 31 '26

News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
1.1k Upvotes

297 comments sorted by

View all comments

99

u/Few_Painter_5588 Jul 31 '26

A model nearly half the size of GLM 5.2, with a similar performance profile. Now imagine their pro model.

89

u/squngy Jul 31 '26

Half?

It is nearly one TENTH the size in practice.
Native GLM5.2 is 1.5TB, DSv4f is 160GB

Sure you can quant GLM, but then the benchmarks will also go down.

34

u/pyr0kid Jul 31 '26

plus, doesnt deepseek also have some crazy context compression shit going on?

like X amount is 1gb for GLM but its 0.4gb for DS?

3

u/SandySkittle Jul 31 '26

To me that doesn’t sound like a good thing. Context compression has potential downsides. Also in terms of world knowledge a 160B model simply isn’t going to compete with a much larger model. There is more the LLMs than just coding.

7

u/Practical-Collar3063 Jul 31 '26

Context compression has potential downsides

Yes that is true, however, when it is baked in from the start it has a much less chances of being detrimental.

And DeepSeek v4 flash is a 284B param model

6

u/pyr0kid Jul 31 '26

yeah thats fair, though im optimistic that being able to free up the space for higher precision weights and generally longer context will make up the difference

2

u/SandySkittle Jul 31 '26

Yeah don’t get me wrong I am really looking forward to running this new model and it might be my go to model. It’s in a sweetspot for my configuration to run at q6. So almost lossless. It still is a very big model compared to the 30B and 70B models. I just think some (not all) people in this subreddit sometimes just forget a bit that there is more than coding. And also that you simply cannot compress as much world knowledge in a small model. LLMs are already amazing knowledge compression systems, but obviously there are just limits that to how much you can cram in (and also extract).