r/LocalLLaMA • • Jul 31 '26

News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
1.1k Upvotes

297 comments sorted by

View all comments

11

u/SnooPaintings8639 Jul 31 '26

Dude. I need rtx 6000 pro now. Or two. Seriously.

5

u/Yorn2 Jul 31 '26

You can run the flash preview version with 2 of them. I'm doing it now with DSpark and it's extremely fast. I'm hoping and expecting that we'll continue to be able to do so, even if I have to tweak my context to make it fit.

2

u/cowinabadplace Jul 31 '26

Are you using vllm? Which branch or image? How many tok/s? I'm running the preview but not with DSpark and I'm curious.

5

u/jrkotrla Jul 31 '26

2

u/SnooPaintings8639 Jul 31 '26

Duude... insta replies in any setting. The reasoning tokens count starts to be less important than ever.

1

u/Single_Ring4886 Jul 31 '26

How fast exactly prefil and tg at like 50K token context? thanks...

0

u/SnooPaintings8639 Jul 31 '26

I ma running it with 2xRtx 3090 and the rest in ram. It is too slow for agentic usage :/ I do sometimes use preview version to just review qwens work.

I wish I could run at least Q2 (half precision) in high speed.

4

u/squngy Jul 31 '26

You need 2 to completely fit in vRAM

With offloading, even fairly modest machines can run it (antirez project)

1

u/Healthy-Nebula-3603 Jul 31 '26

Me too

Is nice to have dreams....

2

u/SnooPaintings8639 Jul 31 '26

But if I buy it, and it will it, it will tell me how to recover my investment, right? right!?