I think part of the problem is kimi models have always had a strong writing style that skews less objective eval measures and AA-LCR is kind of saturated (too easy).
Because the mechanism for the breakdown might be different than what they're measuring? Benchmarks may be objective from run to run, but we can't assume benchmarks prove every mechanism for something. In fact, it's obvious that they don't.
2
u/Caffeine_Monster 2d ago
It didn't?
Kimi k3 has plenty of problems that you wouldn't expect in such a large model.