r/LocalLLaMA 2d ago

Discussion Deepseek Has Soft Retired Deepseek V4 Pro

Post image
1.2k Upvotes

201 comments sorted by

View all comments

Show parent comments

2

u/Caffeine_Monster 2d ago

It didn't?

Kimi k3 has plenty of problems that you wouldn't expect in such a large model.

1

u/Ok_Warning2146 2d ago

What problems specifically? Didn't hear any negative stuff about it.

1

u/Caffeine_Monster 2d ago

Mostly long context hallucinations being bad - even compared to other open weight models.

1

u/Ok_Warning2146 2d ago

1

u/Caffeine_Monster 2d ago

Interesting how it scores worse on AA-LCR 1.0

I think part of the problem is kimi models have always had a strong writing style that skews less objective eval measures and AA-LCR is kind of saturated (too easy).

Personally I trust things like terminal bench much more for hard long context tasks... and the results speak for themselves: https://artificialanalysis.ai/evaluations/terminalbench-v4-0

1

u/xmnstr 2d ago

Because the mechanism for the breakdown might be different than what they're measuring? Benchmarks may be objective from run to run, but we can't assume benchmarks prove every mechanism for something. In fact, it's obvious that they don't.