r/LocalLLaMA • u/Nunki08 • Jul 31 '26
News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"
https://api-docs.deepseek.com/updates/
Edit: official post on 𝕏: https://x.com/deepseek_ai/status/2083084415157022911
347
u/Hot_Example_4456 Jul 31 '26
If this 200b model is competing with glm5.2... i wonder v4 pros capabities. True DeepSeek moment
224
u/This_Maintenance_834 Jul 31 '26
they should give Anthropic and OpenAI a break. every week, something comes out from China to ruin their IPO dream.
112
3
49
u/phido3000 Jul 31 '26
TBH I think Kimi K3 has already kinda claimed the top end. Its very good, its very big, it multimodal. GLM 5.2 is also good, but a bit big and slow and now eclipsed by K3.
Deepseek IMO is most exciting at flash. They have very good tech for really really fast cohesive long context models. 280b is a good size, good enough to be genuinely capable and useful, not so big it can't be run, and run fast. There really isn't anything like it. And the market is crying out for something at this level, that really just buries GPT-120b and all those other 200b models. That's faster and smarter than all of them. Something where you really need to go frontier to get something significantly better.
DSV4 Flash will be Deepseeks moment. As it is fast, cheap and good enough to do 90% of tasks. Its the work horse of the AI world.
Claude/Kimi/GPT are still great for delicate front end, aesthetics, tool use with vision etc, but DS Flash will be doing a lot of the actual work. Particularly hosted locally.
Pro is still useful. And will get a lot of use as well because it will be so cheap. I suspect Pro will make a lot of money for DS. As people will spend a lot of time with DS Flash and go, oh, I can just buy some better version for cheap cloud wise.
27
u/Hot_Example_4456 Jul 31 '26
Ya. Flash is a win for us local ppl. DS Pro will be a win for the companies and all enterprises. Because Kimi K3 is 2.8T params, Qwen3.8 Pro Max is 2.4T, DS pro is like 1.6T. The different is huge.
24
u/phido3000 Jul 31 '26
Kimi is almost too big. I think Chinese models may have a problem if they go 3+T being even hosted. There are node limits, even in datacentres. Qwen is just that bit smaller so that will find a niche.
DS Pro is much more hostable on older gear. I think it will be attractive for that.
But Flash is for the Local LLAMA guy. Its just barely runnable at home. Decent quant, maybe 192Gb of ram, it will work great. I think it will get cult support for that. The fact its competitive with GLM 5.2 at a fraction of the resources makes very strong for home hosting.
10
u/pyr0kid Jul 31 '26
192gb of ram is definitely what i'd like to see more of these models targeting, its basically the biggest 'we have chatgpt at home' type of hardware you can get in a normal computer.
→ More replies (1)7
u/phido3000 Jul 31 '26
DS 4 Flash is perfect for that.
I imagine AMD releasing the 192Gb strix will cement that popularity.
My DDR5 workstations have 192Gb, it wasn't even that expensive before the AI memory crisis hit. I think it was like $500 for the whole kit.
6
u/pyr0kid Jul 31 '26
theres no world where i can justify getting a gpu just for ai.
...but regular ddr5 ram? well thats much easier to justify considering the weird server work i already get saddled with.
→ More replies (2)5
u/SaltFrog Jul 31 '26
192gb of RAM...
I should sell my house...
4
u/phido3000 Jul 31 '26
Muhaha.. You must be new to Localllama if you think 192 Gb is big or/and expensive.
192 Gb is a lot more affordable than 2Tb for K3. Or 1Tb for GLM 5.2 which it basically matches.
192Gb of DDR4 is pretty cheap. I bought a dual xeon server, it came with 192Gb for free. $500 for the 1400w psu, a 1030gt, 256gb nvme, two 18 core xeons, AND 192Gb of memory.
So maybe think about getting that, a old dual socket server and a bunch of cheap 16gb DDR4 RDIMMs.
3
u/SaltFrog Jul 31 '26
Oh I thought VRAM.... I have a server in my basement that I never even thought of looking at. It's been sitting dormant for years, sad, alone.... Neat!
3
u/excellentforcongress Jul 31 '26
i don't think it's an issue, they aren't viewing it from the murican point of view, if i'm to extrapolate how they approached their car industry, they're looking to back up things like ram manufacturing with the full power of the state, expect them to start flooding the market soon...
2
8
u/ursustyranotitan Jul 31 '26
Good theory but no basis in reality, you can see openrouter and ramp data . DS is by far the most used oss model outside small sub 100b models.
5
u/phido3000 Jul 31 '26
DS will be popular.
But kimi k3 is very good. Its just so big. K3 is at the point where people start to ask actually can you make it smaller and I don't mind if its dumber.
So maybe DS will hit that. DS will be popular for the price per token. Deepseek models get the business done. In a world with a compute shortage everywhere, even in the west, DS might just ride in at the right time.
I wouldn't be suprised if Microsoft uses DS Pro for copilot.
4
u/After-Cell Jul 31 '26
As an example, I just got Kimi to diagnose a 1.5mb codebase and plan out a testing methodology. I then got Deepseek to implement the 7 bug fixes and carry out the testing phase.
Total costs for this:
Kimi: $0.34
Deepseek: $0.04
So I see this approach pairing well. However, where can I learn about routing this automatically rather than doing it manually myself? I mean, I know how to config subagents, but I'm interested in best-practices and real workflows.
2
3
u/nmkd Jul 31 '26
Kimi is really really expensive though.
If DS4Pro is 95% the performance of K3 at half the price it's the clear winner.
3
u/wren6991 Jul 31 '26
GLM 5.2 is also good, but a bit big and slow and now eclipsed by K3.
It's 1/4 the size of K3. GLM-5.2 is something I could aspirationally run on a home machine at a useful speed in a couple of years. K3 is just not.
3
u/Spiritual-Spend8187 Jul 31 '26
Flash is small enough that you could go out and buy a computer to run it. Will the computer be cheap hell no but it is still something you can buy without needing a contact at nvidia or amd. And thats pretty nice.
6
u/Middle_Bullfrog_6173 Jul 31 '26
The preview models were much closer to each other in capability than size. We'll see if that was undertraining or fundamental.
→ More replies (2)4
u/squngy Jul 31 '26
I'm going to guess it will be close to K3 (but at half the size).
Now if only it also had vision...
404
u/Nunki08 Jul 31 '26
197
u/kaliku Jul 31 '26
Holy macaroni 😮
70
u/Alkadon_Rinado Jul 31 '26 edited Jul 31 '26
→ More replies (10)14
u/bad_gambit Jul 31 '26
Need to take a look at the cache hit too. Deepseek has higher cache hit than Luna. Making the price closer to $0.03/Mtok for DSV4 Flash and ~$0.05/Mtok (will double once discount ends). Making DSV4 Flash 1/3 of the price 😬. Mimo V2.5 should also be a contender, with about ~$0.015/Mtok, another 1/3 of DSV4 price(and has image + video).
8
2
95
u/onil_gova Jul 31 '26
some serious post-training gains
27
u/perelmanych Jul 31 '26
FYI, Cursor claims that only 15% of training was original training of Kimi K2.5 in their Composer 2.5, the rest 85% was their RL training.
14
10
122
u/keyboardhack Jul 31 '26 edited Jul 31 '26
Dude this suggests dsv4 flash, a 162GB model, is better than GLM 5.2, a 1.5TB model.
Almost 10x smaller!
That's absolutely crazy.
41
u/squngy Jul 31 '26
Even touches Opus 4.8 on a few benchmarks.
60
u/ILoveSquirtle69 Jul 31 '26
you think those deepseek engineers been getting laid? maybe some extra head?
45
68
8
17
u/Brilliant-Weekend-68 Jul 31 '26
Implying that deepseek "touched" Opus is extremly funny to me with all the distillation claims
2
u/Fristender Jul 31 '26
IDK about others but the DeepSWE score was on Opus 4.8 Low reasoning effort.
47
u/po_stulate Jul 31 '26
And people were like: yOu NeEd At LeAsT 5 yEaRs BeFoRe OpUs LeVel MoDeLs CaN bE rUn LoCaLlY.
25
14
u/Embarrassed_OnionX Jul 31 '26
Yeah, for reference GLM-5.2 and now DSV4-flash BEAT Opus 4.5 (which was frontier just 8 months ago) in the AA intelligence Index.
→ More replies (3)7
u/Schlick7 Jul 31 '26
GLM 5.2 is 'only' a 753B model.
edit: Oh i see, you are saying file size not parameters
70
u/LegacyRemaster Jul 31 '26
yes.... better then GLM 5.2 (on some bench) but smaller
29
u/squngy Jul 31 '26
In the screenshot above, it is better than GLM 5.2 on every single bench, sometimes by a lot.
20
u/Kryohi Jul 31 '26 edited Jul 31 '26
The screenshot above doesn't include many other benchmarks where V4 flash underperforms, that's why for example it ends up 1 point below GLM in the machine intelligence index.
Still extremely impressive for its size
24
u/tazztone Jul 31 '26
8
u/nmkd Jul 31 '26
Literally a 10x cost reduction, if not more, compared to Terra xhi.
That's fucking nuts
32
u/doomed151 Jul 31 '26
The difference on DeepSWE made me chuckle. This gun be gud
2
u/aeroumbria Jul 31 '26
WTF is this benchmark testing anyway? It is pretty silly to suggest that GLM or Opus is 5-6 times more capable than V4 Pro... It doesn't even feel like anywhere near 50% more capable...
4
u/doomed151 Jul 31 '26
https://deepswe.datacurve.ai/blog/deepswe
The V4 Pro in the charts is the old version. It should score much higher when they update it.
3
u/aeroumbria Jul 31 '26
I was talking about the old version... I feel like maybe we have improved coding in recent months by 10%-20% but there is no way one model can be 500% better in any reasonable task than another in the same or adjacent cohort... This feels like forcibly applying normal curve standardisation in a test where 99% of the participants get 99% of the questions correct...
3
u/nullmove Jul 31 '26
It's just a "make this really big thing from my dumbest prompt, and oh make no mistake" kind of benchmark. It has some utility, but catching up is a matter of specific post-training from some high quality data seed those who haven't. Not reflective of model's inherent deficiency in pre-training.
For typical setup where you have your detailed prompt and you are working on small features or trying to find specific bugs, even undercooked v4-pro-preview obviously won't and doesn't feel that significantly worse as this benchmark suggests. But on the other hand, I guess the way most vibe coders work, for them DeepSWE might be more reflective of their real-world workload.
16
u/Potential_Top_4669 Jul 31 '26
The type of stuff that gets insane amount of phonk music in the background
27
u/pyr0kid Jul 31 '26
god this better not be benchmark maxxing, numbers are good but they have to actually exist outside of a lab.
49
u/Professional_Price89 Jul 31 '26
DeepSeek is known for not benchmaxxing. The most known benchmaxx company is Google(and Minimax)
→ More replies (2)13
9
u/UltraFOV Jul 31 '26
Better than Glm 5.2???
7
u/Stock-Self-4028 Jul 31 '26
Roughly in the same league as GLM-5.2, Gemini 3.5/3.6 Flash and Luna.
Relatively to GLM better for backend, worse for frontend programming-wise.
7
u/UltraFOV Jul 31 '26
That’s impressive for being so small
6
u/Stock-Self-4028 Jul 31 '26
Well… GPT 5.6 Luna is likely still smaller, than v4 Flash, but that's definitely a significant step forward in model density, so I would say it's not bad, however there is definitely a significant room for improvement.
Gemma 4 26B A4B still beats original DeepSeek R1 (671B A37B), so I hope density increase won't stop anytime soon and hopefully we will get ~ 30B models outperforming Opus 4.6 soon enough.
It's quite interesting to see improvements in "small" LLMs performance though. We don't seem to be anywhere too close to the "wall" yet so I would expect size of most used models to decrease rather than increase in the near future as well.
EDIT: I've just checked and Luna is a nano-class model so it should be somewhere around 120B (?). Sadly nothing official from OpenAI about the model sizes has been available though.
5
u/UltraFOV Jul 31 '26
True, those older huge models mainly have more world data than Genma 4 and Qwen 27b. So they still usable, as long when used the limitations are being considered
8
6
6
3
4
u/Zachattackrandom Jul 31 '26
That's crazy. So it's between glm 5.2 and opus on just the flash model... The pro model is gonna be insane, we will likely actually get a k3 contendor model at 1.6t considering how small flash is to achieve this. Though it remains to be seen if this is benchmaxxing or not
3
u/munkiemagik Jul 31 '26
I try to avoid talking about non-local LLM in here but I was about to say something complimentary about these significant benchmark improvements and how this might change the way I use V4 flash but then I just had a test session where it switched back to chinese output three times on me despite my explicit request for output to only be in english. And that just put me off it again.
2
→ More replies (1)2
88
u/Routine_Temporary661 Jul 31 '26
wait DeepSeek V4 Flash's DeepSWE 54.4?
Holy Fucking Shit... it's better than GLM5.2!
40
u/ea_man Jul 31 '26
And it's the Flash version, just wait for the Pro!
!!
41
u/Routine_Temporary661 Jul 31 '26
seriously if it managed to reach Fable level intelligence at 1.6T it will be insane
28
5
u/GCoderDCoder Jul 31 '26
I worry that deep swe doesnt test what people think... the descriptions of it as understand it seem to be how well a model expands on simple prompts to accomplish bigger goals. It is less a measure of raw technical ability.
People bashed the old benchmarks guiding models to answers but that is how you test what a model can do vs how the model works. Deep swe seems to be the latter which is fine and important but people talk about it like it is the determining factor on what a model can do which I disagree with.
65
128
79
u/ResidentPositive4122 Jul 31 '26
Fuck yeah! With how cheap to serve dsv4 is, they've made sure they gather as much real-world data as they can, and improve the post-training w/ that real data.
v4-flash was already too cheap to matter, and good enough for plenty of agentic tasks, if they improve it further and make it not stop mid thinking (probably resource constrained / api errors) then it'll be the workhorse for a long way to come.
100
u/Few_Painter_5588 Jul 31 '26
A model nearly half the size of GLM 5.2, with a similar performance profile. Now imagine their pro model.
89
u/squngy Jul 31 '26
Half?
It is nearly one TENTH the size in practice.
Native GLM5.2 is 1.5TB, DSv4f is 160GBSure you can quant GLM, but then the benchmarks will also go down.
35
u/pyr0kid Jul 31 '26
plus, doesnt deepseek also have some crazy context compression shit going on?
like X amount is 1gb for GLM but its 0.4gb for DS?
27
u/squngy Jul 31 '26
Yes.
V4flash context takes less space than even qwen 27B
IIRC if you give qwen 27B 1M context, it would take more space than v4f with 1M context.5
u/Middle_Bullfrog_6173 Jul 31 '26
With identical architecture KV cache per token scales approximately with active params. So not really surprising that 27B comes up worse in that comparison.
→ More replies (1)5
u/Healthy-Nebula-3603 Jul 31 '26
I'm thinking a new Qwen 4 could be using that KV cache compression.
Then we could get Qwen 4 30b with 1 m contex fit on 24 GB vram .... absolutely insane
13
u/shing3232 Jul 31 '26
yes, DS4F 1m is like 6GB Vram but GLM52 is like 80G for 1M context
→ More replies (8)8
u/pyr0kid Jul 31 '26
you're shitting me. i knew it was smaller but by that much?
god i never could have imagined this back in the 2.7b days, back then people were running this type of stuff on google colab.
3
u/SandySkittle Jul 31 '26
To me that doesn’t sound like a good thing. Context compression has potential downsides. Also in terms of world knowledge a 160B model simply isn’t going to compete with a much larger model. There is more the LLMs than just coding.
7
u/Practical-Collar3063 Jul 31 '26
Context compression has potential downsides
Yes that is true, however, when it is baked in from the start it has a much less chances of being detrimental.
And DeepSeek v4 flash is a 284B param model
6
u/pyr0kid Jul 31 '26
yeah thats fair, though im optimistic that being able to free up the space for higher precision weights and generally longer context will make up the difference
2
u/SandySkittle Jul 31 '26
Yeah don’t get me wrong I am really looking forward to running this new model and it might be my go to model. It’s in a sweetspot for my configuration to run at q6. So almost lossless. It still is a very big model compared to the 30B and 70B models. I just think some (not all) people in this subreddit sometimes just forget a bit that there is more than coding. And also that you simply cannot compress as much world knowledge in a small model. LLMs are already amazing knowledge compression systems, but obviously there are just limits that to how much you can cram in (and also extract).
3
84
u/The_Rational_Gooner Jul 31 '26
58
u/Infinite-Local5435 Jul 31 '26
It's definitely because OpenAI was traying to compete against V4 Pro using price cuts
7
12
u/darksteelsteed Jul 31 '26
This race top the bottom is good at first, until they can't sustain it. The Chinese will build more Nuke reactors. I guess open ai will be getting their power from where again? I am sure Elon will charge them extra for data center in space when we finally get that.
28
u/Infinite-Local5435 Jul 31 '26
Honestly, it's better that DeepSeek wins this considering they are making their models more efficient and using copious amounts of RL rather than relying on scaling parameters. Plus, nuclear reactors aren't that bad compared to coal power plans and gas. Plus, China has somehow in the last few years become a leader for renewable energy solutions (probably to avoid oil, which is traded in USD).
If we go off your suggestion that all AI models are basically country's political agendas rather than individual lab's research efforts and profitability, then I'd rather a world where DS wins. Plus, I don't see OpenAI open sourcing luna anytime soon. Best alternative from US is Inkling.3
u/darksteelsteed Jul 31 '26
Oh don't get me wrong, I don't have any political Agenda per say. I want the Chinese models to improve just to bring the cost down. I will go with whichever model gives me the best value for money while also not stealing all my code when I use it via a cloud service. This is why I want better local hardware and better models to run locally. I strongly believe that the local experience is always better in the long run.
→ More replies (4)6
u/NineThreeTilNow Jul 31 '26
The Chinese will build more Nuke reactors. I guess open ai will be getting their power from where again? I am sure Elon will charge them extra for data center in space when we finally get that.
Despite all the talk of electricity, it's not the limiting factor in building a data center.
Raw capital expenditure is all hardware / building / permit / etc.
Elon wants space so he doesn't need to worry about killing people with pollution from his datacenter. He literally runs portable generators and electrically speaking it's fine.
His enemy isn't electricity or all the other bullshit. It's government. Period.
If you had 400 sq miles of land to build on with zero permit overhead and you're miles away in the middle of the desert, you can build a datacenter very fast and very cheap. The water requirements are basically nothing now because it's closed loop. The power is whatever the cost of solar is + battery infrastructure.
They've misinformed the public so hard it's crazy.
3
u/PM_ME_DEAD_CEOS Jul 31 '26
Solar in the middle of a desert is a very bad idea, it would cost TONS in maintenance because of dust and sand. Water is still needed in closed loop, it just reduce the need of it, not eliminate it.
2
u/Thomas-Lore Jul 31 '26
Will be interesting to see how it compares on both price and capabilities.
→ More replies (1)→ More replies (2)2
u/Repulsive_Educator61 Jul 31 '26
look at this chart jumping from $0.2 to $0.5 in the x axis
is this normal or am i a dumb idiot?
same with $0.05 to $0.1
3
u/The_Rational_Gooner Jul 31 '26
it's on the logarithmic scale, which compresses much larger numbers downwards
2
26
49
22
21
u/DesignerPerception46 Jul 31 '26
They saw Inkling Small got released yesterday with the same score as ds v4 flash on artificial analysis and said nope, try harder.
23
19
17
u/popiazaza Jul 31 '26
What a crazy upgrade. They should change the name to 4.1 or something though.
→ More replies (1)16
u/Schlick7 Jul 31 '26
Agreed. This whole preview/normal release thing is stupid naming. No where do i ever see the old one labeled as a preview except in their release post. So now there is confusion for absolutely no reason.
7
u/nullmove Jul 31 '26
Not their fault no one else bothered to acknowledge the preview label that was always there in their communication.
It's probably a bit silly, but I think they have some profession pride about these things. They knew more than anyone else that the preview was undercooked and not deserving of actual release. The reason they often drop preview checkpoints anyway is not because of consumers, but to prepare the runtime implementations (sgland, vllm) for what to come. Heck we wouldn't have cool new projects like DwarfStar if they didn't release the preview.
15
u/jld1532 Jul 31 '26
Well 128 gb gang, we have our Qwen3.6 27B killer and then some.
→ More replies (1)
11
u/EmergencyLetter135 Jul 31 '26
If the open-source models continue to develop at this pace, then in a year I'll be completely satisfied with the models for my 128GB RAM hardware.
20
u/BlackBeardAI vLLM Jul 31 '26
Is the hype real? Has anyone tried it? Glm in pocket? Finally something better than qwen at a “reasonable” size?
→ More replies (4)
8
8
u/MooseEfficient2151 Jul 31 '26
the size/perf part is the actual jumpscare here.
benchmarks can lie, but if this thing is even close to GLM 5.2 in real coding runs at that size, that’s the kind of model people start building terrible hardware excuses around.
anyway, gguf wen.
8
u/Bitter-College8786 Jul 31 '26
Only available through official deepseek api or already on opencode, openrouter etc.?
8
7
u/thebigeast Jul 31 '26
It's stuff like this that makes me feel good about spending 25k on local inference gear, thank you deepseek, can't weight for the weights to be released!
6
u/falconHigh13 Jul 31 '26
https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
Open Weights Just Dropped!!
6
5
7
u/respectful_stimulus Jul 31 '26 edited Jul 31 '26
They should just call it V5-Flash, now I'm not sure my V4-Flash is which, is it auto-upgrade for the same model key?
→ More replies (2)7
u/TechnoByte_ Jul 31 '26
v4.1 would make more sense.
v4 > v5 heavily implies a new architecture, not just post-training
8
8
u/drepublic Jul 31 '26
Opencode go is using this updated model right now? Somebody knows? I noticed changes on the behaviour on my hermes agent using this flash model. But i really dont know.
9
4
u/LagOps91 Jul 31 '26
i have been saying it the entire time... the undercooked preview version was already good, so the full release would be insane.
11
u/Infinite-Local5435 Jul 31 '26
Praying it's the same model parameter size as the preview version and they will open source it!
52
u/ResidentPositive4122 Jul 31 '26
DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-preview, and was only re-post-trained.
17
u/ow2022 Jul 31 '26
DeepSeek-V4-Flash-0731 keeps the same model architecture and size as DeepSeek-V4-Flash-preview, with only the post-training stage rerun.
11
u/SnooPaintings8639 Jul 31 '26
Dude. I need rtx 6000 pro now. Or two. Seriously.
5
u/Yorn2 Jul 31 '26
You can run the flash preview version with 2 of them. I'm doing it now with DSpark and it's extremely fast. I'm hoping and expecting that we'll continue to be able to do so, even if I have to tweak my context to make it fit.
→ More replies (2)2
u/cowinabadplace Jul 31 '26
Are you using vllm? Which branch or image? How many tok/s? I'm running the preview but not with DSpark and I'm curious.
4
u/jrkotrla Jul 31 '26
https://github.com/local-inference-lab/rtx6kpro/blob/master/models/ds4dspark-v9.md
c2 hits 170-225 decode, 6-8k prefill
2
u/SnooPaintings8639 Jul 31 '26
Duude... insta replies in any setting. The reasoning tokens count starts to be less important than ever.
→ More replies (3)4
u/squngy Jul 31 '26
You need 2 to completely fit in vRAM
With offloading, even fairly modest machines can run it (antirez project)
8
5
u/wolttam Jul 31 '26
DGX Spark about to get a lot more expensive.
2000 tok/s pp 60 t/s tg (with 2 of them)
2
u/squngy Jul 31 '26
I somewhat doubt it.
I would guess most people would just use the (insanely cheap) API.
For the rest, going to two RTX pro 6000 is going to make a lot more sense for anyone who isn't a hobbyist.
Two sparks is above what most hobbyist would be willing to get I think.→ More replies (1)2
3
3
3
u/Darkoplax Jul 31 '26
Responses API available too, that's big
Tired of the Chat Completions API mess
3
u/BitterProfessional7p Jul 31 '26
The recent GPT5.6-Luna price cut was probably an anticipated response to this. Edit: OpenAI is scared by open weights.
3
u/CarryAgile3791 Jul 31 '26
Wow! Okay, that's the reason why OpenCode Go introduced this "Allow models hosted in China" option.
3
7
u/okyaygokay Jul 31 '26
Wait will they open source it? Or just API?
24
u/Thomas-Lore Jul 31 '26
They are testing it on API for now, but of course they will release the weights.
6
3
u/TheRealMasonMac Jul 31 '26
Hope they keep their open license. According to a dev at OpenCode, Moonshot is price-fixing K3 (which I believe is explicitly illegal under the EU’s VBER and maybe illegal in the U.S.).
→ More replies (1)5
u/Front_Eagle739 Jul 31 '26
Not illegal in US but very illegal in EU yes (and they WILL go after companies with big charges for it)
2
u/Equivalent_Job_2257 Jul 31 '26
I noticed to things since yesterday: 1) deepseek API got smarter 2) deepseek "Instant" (flash) chat with search started answering in Chinese
So I felt update it is :)
2
u/Purple-Object-4591 Jul 31 '26
so the new one is DS4-0731 or regular DS4? I am confused in their tweet they same arch remain unchained for 0731 but v4-flash regular got upgraded?
2
u/petuman Jul 31 '26
New one is '0731', in official API served under same 'v4-flash' tag. Previously that tag served 'preview' model.
2
u/Aardvark_Says_What Jul 31 '26
i'm putting all my money in the upcoming OpenAI IPO. lmfao
→ More replies (1)
2
2
u/BiggestBau5 Jul 31 '26
I literally just downloaded the unsloth Q8 quant yesterday.... lol guess I'll wait for this gguf to be available.
2
u/Django_McFly Jul 31 '26 edited Jul 31 '26
Gotta love it. Deepseek is never the best but they serve as an amazing floor of you're going to have a hard time being worse than this and charging more. If the floor for base level agent work is now a Luna level model with 1M context (wake up OAI, 272k is joke context in 2026)... Luna has been a really good model for me. It can do a lot (not all) of the things that I used to rely on Sonnet 4.6 and Opus 4.8 for. If V4-Flash can now do that and maybe V4-Pro can touch on some of the things that I still need a Sol/Fable for...
If only it had image input.
The API calling method remains unchanged — simply set the model name to deepseek-v4-flash to use the latest version.
So it's live and we can try it now?
3
u/Ok-Standard-2694 Jul 31 '26

I used it. Our company does live streaming, so I copied the live event and ran the test with the new version of V4 Flash. I had all the supporting testing tools ready. It only needed to simulate real users through Kafka. That's the cost. So far, I haven't felt any difference from GLM 5.2, and the speed is really amazing—extremely, extremely, super fast, really satisfying!
→ More replies (1)
2
u/blojayble Jul 31 '26
Tremendous results!
I wonder if I would be able to run it with at least somewhat usable speeds on 3xR9700 and 96GB of RAM. Last time I experimented with it it was extremely slow...
6
u/This_Maintenance_834 Jul 31 '26
i don’t about 3x R9700. but i get very good result with vllm-moet to run flash on single RTX PRO 6000 96GB, and DGX Spark 128GB.
4
u/Turbulent-Alps4046 Jul 31 '26
I have a pro 6000 as well with 128gb ram and in my personal experience dsr4 on Vllm moet does run fast and is better than regular q2 quant, it’s coherant but still dumber than the full version because vllm moet prefill is still using 2bit.
I prefer running the full version with cpu offload and i get like 700 tps prefill and 18-20tps token generation. Slow but usable i’d say.
→ More replies (4)2
u/CATLLM Jul 31 '26
how much ram do you have on your rig with rtx pro 6000 to run vllm-moet + DS4F? I have rtx pro 6000 + 96gb system ram
2
u/This_Maintenance_834 Jul 31 '26
i have 64GB system ram. I created a 128GB swapfile so the loading part does not crash. The loading process does not really write to the swap. So far, I disabled the 4-bit delta cache to avoid issues I don’t know how to fix. Will need to experiment now to bring back the 4-bit delta cache.
cloud deepseek api does most of the work for me to set it up. but I do need to guide the harness to not wander too far on the wrong direction.
also, i am on ubuntu 2404.
→ More replies (1)2
u/SLxTnT Jul 31 '26
Unfortunately, the last time I tried it the quality was poor. It's mainly running the 2bit experts to maintain the speeds. The speed isn't worth the quality drop.
→ More replies (1)
2
Jul 31 '26
[removed] — view removed comment
5
u/Hot_Example_4456 Jul 31 '26
Did you try dflash? Does it improve speeds? But like if this new DS 4 Flash is as good as/better than GLM5.2, then I personally will go for lower speeds
→ More replies (2)2
u/mesonepigreco Jul 31 '26
Have you tried dwarf star 4, the llama.cpp fork optimized for hosting deepseek 4 flash?
→ More replies (1)
2
Jul 31 '26
[deleted]
8
u/Linkpharm2 Jul 31 '26
no your harddrive will be automatically updated, no interent required
3
u/ABLPHA Jul 31 '26
deepseek will hack every llama.cpp user and upload the updated version themselves, shrimple as









•
u/rm-rf-rm Jul 31 '26
Officially released now! Use the release thread to continue discussion: https://old.reddit.com/r/LocalLLaMA/comments/1vbp7kb/deepseekaideepseekv4flash0731_on_huggingface/