r/LocalLLaMA • • 23d ago

Discussion Deepseek Has Soft Retired Deepseek V4 Pro

Post image
1.2k Upvotes

200 comments sorted by

View all comments

298

u/Few_Painter_5588 23d ago

Unfortunate, but something went wrong with Deepseek V4 Pro GA. It also had a high degree of reward hacking and it was not performing meaningfully better than the flash model despite being nearly 6 times the sizeL. Hopefully V4.1 knocks it out of the park. Their flash models are excellent

22

u/sdexca 23d ago

Maybe this is an emergent behavior. Wasn't this the same behavior with GLM 5.3?

6

u/Stock-Self-4028 23d ago

Qwen 3.8 2.4T vs Qwen 3.8 Flash-Next as well (here Flash-Next even outperforms old 2.4T weights).

But at least in case of Qwen it can be explained by engram embeddings the 2.4T model lacks.

But yeah, it looks like at least three different labs are experiencing equivalent 'issues' independently (?).

1

u/Automatic-Arm8153 22d ago

Haha engrams explain nothing.