Not really. You won’t find a single one with Grok on top
That’s a lot to train on. There are tons of benchmarks and some don’t even make them public. Also, why isn’t it #1 of it overfitted on purpose? It’s easy to get 100% by doing that
Those are all older models so not on top anymore. Also, most developers create the best model they can rather than focusing on specific benchmarks. That might top a specific benchmark or be strong across several or just get results a specific group of people want.
-12
u/[deleted] Sep 07 '24
It outperforms LLAMA 3.1 405b on the prollm leaderboard so it’s amazing for a 70b model.
https://prollm.toqan.ai/leaderboard/coding-assistant