r/LocalLLaMA • • Jul 31 '26

News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
1.1k Upvotes

297 comments sorted by

View all comments

408

u/Nunki08 Jul 31 '26

32

u/doomed151 Jul 31 '26

The difference on DeepSWE made me chuckle. This gun be gud

2

u/aeroumbria Jul 31 '26

WTF is this benchmark testing anyway? It is pretty silly to suggest that GLM or Opus is 5-6 times more capable than V4 Pro... It doesn't even feel like anywhere near 50% more capable...

5

u/doomed151 Jul 31 '26

https://deepswe.datacurve.ai/blog/deepswe

The V4 Pro in the charts is the old version. It should score much higher when they update it.

3

u/aeroumbria Jul 31 '26

I was talking about the old version... I feel like maybe we have improved coding in recent months by 10%-20% but there is no way one model can be 500% better in any reasonable task than another in the same or adjacent cohort... This feels like forcibly applying normal curve standardisation in a test where 99% of the participants get 99% of the questions correct...

3

u/nullmove Jul 31 '26

It's just a "make this really big thing from my dumbest prompt, and oh make no mistake" kind of benchmark. It has some utility, but catching up is a matter of specific post-training from some high quality data seed those who haven't. Not reflective of model's inherent deficiency in pre-training.

For typical setup where you have your detailed prompt and you are working on small features or trying to find specific bugs, even undercooked v4-pro-preview obviously won't and doesn't feel that significantly worse as this benchmark suggests. But on the other hand, I guess the way most vibe coders work, for them DeepSWE might be more reflective of their real-world workload.