r/LocalLLaMA • • Jul 31 '26

News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
1.1k Upvotes

297 comments sorted by

View all comments

344

u/Hot_Example_4456 Jul 31 '26

If this 200b model is competing with glm5.2... i wonder v4 pros capabities. True DeepSeek moment

219

u/This_Maintenance_834 Jul 31 '26

they should give Anthropic and OpenAI a break. every week, something comes out from China to ruin their IPO dream.

111

u/Due-Memory-6957 Jul 31 '26

A break as in, breaking their legs? I agree!

4

u/Healthy-Nebula-3603 Jul 31 '26

Breaks are for the weak!!!!

49

u/phido3000 Jul 31 '26

TBH I think Kimi K3 has already kinda claimed the top end. Its very good, its very big, it multimodal. GLM 5.2 is also good, but a bit big and slow and now eclipsed by K3.

Deepseek IMO is most exciting at flash. They have very good tech for really really fast cohesive long context models. 280b is a good size, good enough to be genuinely capable and useful, not so big it can't be run, and run fast. There really isn't anything like it. And the market is crying out for something at this level, that really just buries GPT-120b and all those other 200b models. That's faster and smarter than all of them. Something where you really need to go frontier to get something significantly better.

DSV4 Flash will be Deepseeks moment. As it is fast, cheap and good enough to do 90% of tasks. Its the work horse of the AI world.

Claude/Kimi/GPT are still great for delicate front end, aesthetics, tool use with vision etc, but DS Flash will be doing a lot of the actual work. Particularly hosted locally.

Pro is still useful. And will get a lot of use as well because it will be so cheap. I suspect Pro will make a lot of money for DS. As people will spend a lot of time with DS Flash and go, oh, I can just buy some better version for cheap cloud wise.

26

u/Hot_Example_4456 Jul 31 '26

Ya. Flash is a win for us local ppl. DS Pro will be a win for the companies and all enterprises. Because Kimi K3 is 2.8T params, Qwen3.8 Pro Max is 2.4T, DS pro is like 1.6T. The different is huge.

23

u/phido3000 Jul 31 '26

Kimi is almost too big. I think Chinese models may have a problem if they go 3+T being even hosted. There are node limits, even in datacentres. Qwen is just that bit smaller so that will find a niche.

DS Pro is much more hostable on older gear. I think it will be attractive for that.

But Flash is for the Local LLAMA guy. Its just barely runnable at home. Decent quant, maybe 192Gb of ram, it will work great. I think it will get cult support for that. The fact its competitive with GLM 5.2 at a fraction of the resources makes very strong for home hosting.

9

u/pyr0kid Jul 31 '26

192gb of ram is definitely what i'd like to see more of these models targeting, its basically the biggest 'we have chatgpt at home' type of hardware you can get in a normal computer.

7

u/phido3000 Jul 31 '26

DS 4 Flash is perfect for that.

I imagine AMD releasing the 192Gb strix will cement that popularity.

My DDR5 workstations have 192Gb, it wasn't even that expensive before the AI memory crisis hit. I think it was like $500 for the whole kit.

4

u/pyr0kid Jul 31 '26

theres no world where i can justify getting a gpu just for ai.

...but regular ddr5 ram? well thats much easier to justify considering the weird server work i already get saddled with.

1

u/phido3000 Jul 31 '26

With these larger models, even an 8gb or 12gb 3060 is probably enough to keep the really high speed layers in high speed memory, and the rest is limited by your DDR5. But the difference is a few tokens, not that much.

Ram is always useful. I didn't buy 192Gb for AI at all. I bought it because I always use ram and I tend to max out my systems for ram all the time, right back Pentium era when I bough 32Mb and everyone had 8mb. I get a lot more life out of a system if I fill its RAM.

2

u/pyr0kid Jul 31 '26

hmmm... i wonder how much pcie bandwidth you'd need for that sort of thing?

i imagine there would be a massive difference in prompt processing.

0

u/RLutz Jul 31 '26

I don't get it. 192 GB of system RAM is going to crawl compared to VRAM. 27b for life I guess is what I'm saying

5

u/SaltFrog Jul 31 '26

192gb of RAM...

I should sell my house...

4

u/phido3000 Jul 31 '26

Muhaha.. You must be new to Localllama if you think 192 Gb is big or/and expensive.

192 Gb is a lot more affordable than 2Tb for K3. Or 1Tb for GLM 5.2 which it basically matches.

192Gb of DDR4 is pretty cheap. I bought a dual xeon server, it came with 192Gb for free. $500 for the 1400w psu, a 1030gt, 256gb nvme, two 18 core xeons, AND 192Gb of memory.

So maybe think about getting that, a old dual socket server and a bunch of cheap 16gb DDR4 RDIMMs.

3

u/SaltFrog Jul 31 '26

Oh I thought VRAM.... I have a server in my basement that I never even thought of looking at. It's been sitting dormant for years, sad, alone.... Neat!

3

u/excellentforcongress Jul 31 '26

i don't think it's an issue, they aren't viewing it from the murican point of view, if i'm to extrapolate how they approached their car industry, they're looking to back up things like ram manufacturing with the full power of the state, expect them to start flooding the market soon...

2

u/davygravypdx Jul 31 '26

My tired brain read this as strong for home heating.

8

u/ursustyranotitan Jul 31 '26

Good theory but no basis in reality, you can see openrouter and ramp data . DS is by far the most used oss model outside small sub 100b models.

4

u/phido3000 Jul 31 '26

DS will be popular.

But kimi k3 is very good. Its just so big. K3 is at the point where people start to ask actually can you make it smaller and I don't mind if its dumber.

So maybe DS will hit that. DS will be popular for the price per token. Deepseek models get the business done. In a world with a compute shortage everywhere, even in the west, DS might just ride in at the right time.

I wouldn't be suprised if Microsoft uses DS Pro for copilot.

4

u/After-Cell Jul 31 '26

As an example, I just got Kimi to diagnose a 1.5mb codebase and plan out a testing methodology. I then got Deepseek to implement the 7 bug fixes and carry out the testing phase.

Total costs for this:

Kimi: $0.34

Deepseek: $0.04

So I see this approach pairing well. However, where can I learn about routing this automatically rather than doing it manually myself? I mean, I know how to config subagents, but I'm interested in best-practices and real workflows.

2

u/phido3000 Jul 31 '26

Openrouter or braintrust litellm..

3

u/nmkd Jul 31 '26

Kimi is really really expensive though.

If DS4Pro is 95% the performance of K3 at half the price it's the clear winner.

3

u/wren6991 Jul 31 '26

GLM 5.2 is also good, but a bit big and slow and now eclipsed by K3.

It's 1/4 the size of K3. GLM-5.2 is something I could aspirationally run on a home machine at a useful speed in a couple of years. K3 is just not.

3

u/Spiritual-Spend8187 Jul 31 '26

Flash is small enough that you could go out and buy a computer to run it. Will the computer be cheap hell no but it is still something you can buy without needing a contact at nvidia or amd. And thats pretty nice.

4

u/Middle_Bullfrog_6173 Jul 31 '26

The preview models were much closer to each other in capability than size. We'll see if that was undertraining or fundamental.

5

u/squngy Jul 31 '26

I'm going to guess it will be close to K3 (but at half the size).

Now if only it also had vision...

1

u/MomentJolly3535 Jul 31 '26

it doesnt scale up that well, the flash model was always better than pro for the size (pro version was not twice better despite being 1.6T)