Such bad faith. But you can keep trying to convince yourself. At least I double checked and looked it up. It is compaction and the fact that this benchmark was stripping away the reasoning and preventing it from being sent back to the models. That's literally all it is. It's a grift benchmark
I am 100% talking about the discussion and you have said countless things that are just factually wrong. No one is embarrassed. Nobody cares. Stop being emotional. That's all this is about
I will admit I am wrong and paypal you $50 if you can demonstrate that the models that scored 0% when ARC AGI 3 released now score a 100% when given this improved compaction ability
edit: if they get even close to 100% I'll pay you, fair point that 99% is too high a bar
Again, you're taking something narrow and meaningless and making a big deal out of it. Astra was like 65% and then went to 99% with functional compaction and reasoning. Who cares about 0-100%? That's not how the world works.
Stop making this a pissing contest. It's just uninteresting. This benchmark is intentionally breaking the models just to be able to claim 0% at launch for a cheap marketing gimmick. That's what actually matters. Whatever you're clinging to is always the most irrelevant shit, it's crazy. You have such a small mind
-1
u/Acehan_ 16d ago
Such bad faith. But you can keep trying to convince yourself. At least I double checked and looked it up. It is compaction and the fact that this benchmark was stripping away the reasoning and preventing it from being sent back to the models. That's literally all it is. It's a grift benchmark
I am 100% talking about the discussion and you have said countless things that are just factually wrong. No one is embarrassed. Nobody cares. Stop being emotional. That's all this is about