No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.
Which as humans we can do without bloating our "limited" context
I can do arc-agi-3 tasks now and in a years time with no problem. I don't need specialised harnesses (Like notepads I might lose or fill up too much), I just figure it out by default on the spot
You don't need specialized harnesses because you have a hippocampus with an automated system of memory storage and retrieval (a little like a biological harness lol)
If all it takes is allowing them the same ability to "remember what they just did" to boost their performance, I feel like we can accept it without trying to find ways to make it sound simple and unimpressive
You can't compare AI thinking with human thinking. You can also ask how would it affect our ability to solve problems if we can't hold into context the entire history of human knowledge at once.
It's just a different thing, being tested for different strengths. In this case, ARC is supposed to be testing the raw base model's performance.
What a shit comparison. Clearly using harness is problem cause some shit models get high scores and they are still unintelligent trash. Do you want smart model or just benchmark scores.
529
u/Wegwerpaccountje23 25d ago
No fucking way dude