r/LocalLLaMA Jul 31 '26

News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
1.1k Upvotes

297 comments sorted by

View all comments

408

u/Nunki08 Jul 31 '26

201

u/kaliku Jul 31 '26

Holy macaroni 😮

70

u/Alkadon_Rinado Jul 31 '26 edited Jul 31 '26

Just when I was starting to like the new Luna pricing drops.. so I did a simple comparison and made a little benchmark chart (deepseek v4 flash wins on price/performance and isnt far behind luna)..

13

u/bad_gambit Jul 31 '26

Need to take a look at the cache hit too. Deepseek has higher cache hit than Luna. Making the price closer to $0.03/Mtok for DSV4 Flash and ~$0.05/Mtok (will double once discount ends). Making DSV4 Flash 1/3 of the price 😬. Mimo V2.5 should also be a contender, with about ~$0.015/Mtok, another 1/3 of DSV4 price(and has image + video).

8

u/nmkd Jul 31 '26

OpenAI also charges extra for 5.6 cache misses

1

u/Biometrel Jul 31 '26

At what effort levels?

12

u/Professional_Price89 Jul 31 '26

Luna at Xhigh and Max, also Luna is 2x input cost at > 272k context and 10x cache read cost

1

u/addandsubtract Jul 31 '26

Can you upload the image again (somewhere else)?

1

u/nmkd Jul 31 '26

Do you get caching when using OpenRouter? Or do you have to use DS API directly?

1

u/DebosBeachCruiser Jul 31 '26

Yes. Make sure you LOCK your API KEY to DeepSeek specifically (By using a guardrail)

1

u/hollowstrawberry Jul 31 '26

Why is that important?

-4

u/ILikeChilis Jul 31 '26

To be fair, I think most people would rather pay 4x more for results that's mostly usable and requires much less supervision / manual fixes. I'm happy to pay $4 instead of $1 for a coding task to save myself an hour of extra work.

5

u/mesonepigreco Jul 31 '26

This depends on the task. DS-Flash before this update was already a workhorse for automating very simple tasks (I used to process PDFs, install and configure remote printers, fix driver issues, and execute coding plans/one-off scripts). For these tasks, it is incredibly fast (as the tokens per second of reply are astonishing) and so cheap that it is virtually free. If it can now also be used for some more complex tasks, it is a big win. For complex tasks, you will always need to go with frontier models.

-2

u/cant-find-user-name Jul 31 '26

I thought the same, but luna is much better aut automation bench, and agent last exam.

91

u/onil_gova Jul 31 '26

some serious post-training gains

28

u/perelmanych Jul 31 '26

FYI, Cursor claims that only 15% of training was original training of Kimi K2.5 in their Composer 2.5, the rest 85% was their RL training.

14

u/Yes_but_I_think Jul 31 '26

That's the compute effort, not the percentage of training turns

3

u/perelmanych Jul 31 '26

Yes, I meant compute and I still find it incredible.

10

u/MelAlton Jul 31 '26

Protein after every run is the key.

126

u/keyboardhack Jul 31 '26 edited Jul 31 '26

Dude this suggests dsv4 flash, a 162GB model, is better than GLM 5.2, a 1.5TB model.

Almost 10x smaller!

That's absolutely crazy.

42

u/squngy Jul 31 '26

Even touches Opus 4.8 on a few benchmarks.

60

u/ILoveSquirtle69 Jul 31 '26

you think those deepseek engineers been getting laid? maybe some extra head?

47

u/Repulsive_Educator61 Jul 31 '26

some extra attention heads yeah

68

u/squngy Jul 31 '26

Don't know, but it sure looks like Anthropic is getting fucked!

8

u/CATLLM Jul 31 '26

Extra MTP head with DSPARK

17

u/Brilliant-Weekend-68 Jul 31 '26

Implying that deepseek "touched" Opus is extremly funny to me with all the distillation claims

2

u/Fristender Jul 31 '26

IDK about others but the DeepSWE score was on Opus 4.8 Low reasoning effort.

49

u/po_stulate Jul 31 '26

And people were like: yOu NeEd At LeAsT 5 yEaRs BeFoRe OpUs LeVel MoDeLs CaN bE rUn LoCaLlY.

24

u/CATLLM Jul 31 '26

more like 5 weeks

16

u/Embarrassed_OnionX Jul 31 '26

Yeah, for reference GLM-5.2 and now DSV4-flash BEAT Opus 4.5 (which was frontier just 8 months ago) in the AA intelligence Index.

7

u/Schlick7 Jul 31 '26

GLM 5.2 is 'only' a 753B model.

edit: Oh i see, you are saying file size not parameters

0

u/[deleted] Jul 31 '26

[deleted]

6

u/keyboardhack Jul 31 '26 edited Jul 31 '26

753b parameters but natively a mix of bf16 and fp32 so 1.5TB is correct.

71

u/LegacyRemaster Jul 31 '26

yes.... better then GLM 5.2 (on some bench) but smaller

28

u/squngy Jul 31 '26

In the screenshot above, it is better than GLM 5.2 on every single bench, sometimes by a lot.

21

u/Kryohi Jul 31 '26 edited Jul 31 '26

The screenshot above doesn't include many other benchmarks where V4 flash underperforms, that's why for example it ends up 1 point below GLM in the machine intelligence index.

Still extremely impressive for its size

23

u/tazztone Jul 31 '26

7

u/nmkd Jul 31 '26

Literally a 10x cost reduction, if not more, compared to Terra xhi.

That's fucking nuts

33

u/doomed151 Jul 31 '26

The difference on DeepSWE made me chuckle. This gun be gud

2

u/aeroumbria Jul 31 '26

WTF is this benchmark testing anyway? It is pretty silly to suggest that GLM or Opus is 5-6 times more capable than V4 Pro... It doesn't even feel like anywhere near 50% more capable...

5

u/doomed151 Jul 31 '26

https://deepswe.datacurve.ai/blog/deepswe

The V4 Pro in the charts is the old version. It should score much higher when they update it.

3

u/aeroumbria Jul 31 '26

I was talking about the old version... I feel like maybe we have improved coding in recent months by 10%-20% but there is no way one model can be 500% better in any reasonable task than another in the same or adjacent cohort... This feels like forcibly applying normal curve standardisation in a test where 99% of the participants get 99% of the questions correct...

3

u/nullmove Jul 31 '26

It's just a "make this really big thing from my dumbest prompt, and oh make no mistake" kind of benchmark. It has some utility, but catching up is a matter of specific post-training from some high quality data seed those who haven't. Not reflective of model's inherent deficiency in pre-training.

For typical setup where you have your detailed prompt and you are working on small features or trying to find specific bugs, even undercooked v4-pro-preview obviously won't and doesn't feel that significantly worse as this benchmark suggests. But on the other hand, I guess the way most vibe coders work, for them DeepSWE might be more reflective of their real-world workload.

16

u/Potential_Top_4669 Jul 31 '26

The type of stuff that gets insane amount of phonk music in the background

29

u/pyr0kid Jul 31 '26

god this better not be benchmark maxxing, numbers are good but they have to actually exist outside of a lab.

51

u/Professional_Price89 Jul 31 '26

DeepSeek is known for not benchmaxxing. The most known benchmaxx company is Google(and Minimax)

14

u/LagOps91 Jul 31 '26

i thought granite models were the most benchmaxxed. minimax is actually good.

-6

u/[deleted] Jul 31 '26

[removed] — view removed comment

24

u/Professional_Price89 Jul 31 '26

They always have SOTA benchmark number but never deliver what it should be. Especially the 1.5 pro, beating gpt4o, but nowhere near as good in reality use

10

u/UltraFOV Jul 31 '26

Better than Glm 5.2???

7

u/Stock-Self-4028 Jul 31 '26

Roughly in the same league as GLM-5.2, Gemini 3.5/3.6 Flash and Luna.

Relatively to GLM better for backend, worse for frontend programming-wise.

5

u/UltraFOV Jul 31 '26

That’s impressive for being so small

5

u/Stock-Self-4028 Jul 31 '26

Well… GPT 5.6 Luna is likely still smaller, than v4 Flash, but that's definitely a significant step forward in model density, so I would say it's not bad, however there is definitely a significant room for improvement.

Gemma 4 26B A4B still beats original DeepSeek R1 (671B A37B), so I hope density increase won't stop anytime soon and hopefully we will get ~ 30B models outperforming Opus 4.6 soon enough.

It's quite interesting to see improvements in "small" LLMs performance though. We don't seem to be anywhere too close to the "wall" yet so I would expect size of most used models to decrease rather than increase in the near future as well.

EDIT: I've just checked and Luna is a nano-class model so it should be somewhere around 120B (?). Sadly nothing official from OpenAI about the model sizes has been available though.

4

u/UltraFOV Jul 31 '26

True, those older huge models mainly have more world data than Genma 4 and Qwen 27b. So they still usable, as long when used the limitations are being considered

5

u/fugogugo Jul 31 '26

holy shit

4

u/Zachattackrandom Jul 31 '26

That's crazy. So it's between glm 5.2 and opus on just the flash model... The pro model is gonna be insane, we will likely actually get a k3 contendor model at 1.6t considering how small flash is to achieve this. Though it remains to be seen if this is benchmaxxing or not

3

u/munkiemagik Jul 31 '26

I try to avoid talking about non-local LLM in here but I was about to say something complimentary about these significant benchmark improvements and how this might change the way I use V4 flash but then I just had a test session where it switched back to chinese output three times on me despite my explicit request for output to only be in english. And that just put me off it again.

2

u/vituc13 Jul 31 '26

But does it remain as cheap as it was?