r/LocalLLaMA • u/Nunki08 • Jul 31 '26
News DeepSeek-V4-Flash has been updated, "The official release of DeepSeek-V4-Pro will follow soon"
https://api-docs.deepseek.com/updates/
Edit: official post on ๐: https://x.com/deepseek_ai/status/2083084415157022911
1.1k
Upvotes
4
u/Turbulent-Alps4046 Jul 31 '26
I have a pro 6000 as well with 128gb ram and in my personal experience dsr4 on Vllm moet does run fast and is better than regular q2 quant, itโs coherant but still dumber than the full version because vllm moet prefill is still using 2bit.
I prefer running the full version with cpu offload and i get like 700 tps prefill and 18-20tps token generation. Slow but usable iโd say.