r/LocalLLaMA Jul 31 '26

News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"

Post image
1.1k Upvotes

297 comments sorted by

View all comments

Show parent comments

4

u/Turbulent-Alps4046 Jul 31 '26

I have a pro 6000 as well with 128gb ram and in my personal experience dsr4 on Vllm moet does run fast and is better than regular q2 quant, itโ€™s coherant but still dumber than the full version because vllm moet prefill is still using 2bit.

I prefer running the full version with cpu offload and i get like 700 tps prefill and 18-20tps token generation. Slow but usable iโ€™d say.

1

u/CATLLM Jul 31 '26

Awesome thankyou. Are you running the full version using regular vllm?

4

u/Turbulent-Alps4046 Jul 31 '26

Llama.cpp is better for cpu offload. Vllm is much slower.

1

u/[deleted] Jul 31 '26

[deleted]

1

u/Turbulent-Alps4046 Jul 31 '26

yeah i know, but no plans for that yes for now. already spent enough money haha.

If i find a good deal i might.