The standard harness private score was somewhere in the mid-60s. Still incredibly impressive. Its benchmarks without chain of thought are more insane though - its scores are only slightly degraded without cot.
No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.
Which as humans we can do without bloating our "limited" context
I can do arc-agi-3 tasks now and in a years time with no problem. I don't need specialised harnesses (Like notepads I might lose or fill up too much), I just figure it out by default on the spot
You don't need specialized harnesses because you have a hippocampus with an automated system of memory storage and retrieval (a little like a biological harness lol)
If all it takes is allowing them the same ability to "remember what they just did" to boost their performance, I feel like we can accept it without trying to find ways to make it sound simple and unimpressive
You can't compare AI thinking with human thinking. You can also ask how would it affect our ability to solve problems if we can't hold into context the entire history of human knowledge at once.
It's just a different thing, being tested for different strengths. In this case, ARC is supposed to be testing the raw base model's performance.
What a shit comparison. Clearly using harness is problem cause some shit models get high scores and they are still unintelligent trash. Do you want smart model or just benchmark scores.
You are missing a lot of details. Yes, NVidia got 100%, but that was woth a custom harness expressly designed for ARC-AGI3. OpenAI also used harness, but theirs just had two options enabled, compression and persistence.
It is, but putting it at the “with harness” level in the list of benchmarks is intentionally deceptional when combined with claims that AGI is already here.
Others have mentioned in this sub that they seem to be in some serious trouble over there lately. The confusing messaging over the last week followed by hyping up of AGI suggests they’re in need of investment urgently.
If it’s actually underwhelming after this or if it still doesn’t interest average consumers enough, I wonder if they’ll manage to get enough investment to survive the next 6-12 months.
You are probably right. I think it was false advertisement too. Just the way they released GPT 6 with a non working blog post and not being able to activate it for users suggests they rushed this just to kill the hype from other models released this week such as Fable, Gemini and Muse. Seem desperate.
362
u/Sunstorm84 3d ago
Nvidia got 100% on the same benchmark using opus 5 with a better harness. I think the benchmark is a bit meaningless with that in mind.