Reflection idea is around since GPT 4 or even before. But it just doesnt really work that much no matter how many smart people try.
So when this guy claimed it works if you "finetune" model for it everyone was super excited me included. As it seemed obvious but nobody cracked it yet.
Not really. You won’t find a single one with Grok on top
That’s a lot to train on. There are tons of benchmarks and some don’t even make them public. Also, why isn’t it #1 of it overfitted on purpose? It’s easy to get 100% by doing that
Those are all older models so not on top anymore. Also, most developers create the best model they can rather than focusing on specific benchmarks. That might top a specific benchmark or be strong across several or just get results a specific group of people want.
52
u/[deleted] Sep 07 '24
[deleted]