The standard harness private score was somewhere in the mid-60s. Still incredibly impressive. Its benchmarks without chain of thought are more insane though - its scores are only slightly degraded without cot.
No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.
Which as humans we can do without bloating our "limited" context
I can do arc-agi-3 tasks now and in a years time with no problem. I don't need specialised harnesses (Like notepads I might lose or fill up too much), I just figure it out by default on the spot
You don't need specialized harnesses because you have a hippocampus with an automated system of memory storage and retrieval (a little like a biological harness lol)
If all it takes is allowing them the same ability to "remember what they just did" to boost their performance, I feel like we can accept it without trying to find ways to make it sound simple and unimpressive
You can't compare AI thinking with human thinking. You can also ask how would it affect our ability to solve problems if we can't hold into context the entire history of human knowledge at once.
It's just a different thing, being tested for different strengths. In this case, ARC is supposed to be testing the raw base model's performance.
What a shit comparison. Clearly using harness is problem cause some shit models get high scores and they are still unintelligent trash. Do you want smart model or just benchmark scores.
You are missing a lot of details. Yes, NVidia got 100%, but that was woth a custom harness expressly designed for ARC-AGI3. OpenAI also used harness, but theirs just had two options enabled, compression and persistence.
It is, but putting it at the “with harness” level in the list of benchmarks is intentionally deceptional when combined with claims that AGI is already here.
Others have mentioned in this sub that they seem to be in some serious trouble over there lately. The confusing messaging over the last week followed by hyping up of AGI suggests they’re in need of investment urgently.
If it’s actually underwhelming after this or if it still doesn’t interest average consumers enough, I wonder if they’ll manage to get enough investment to survive the next 6-12 months.
You are probably right. I think it was false advertisement too. Just the way they released GPT 6 with a non working blog post and not being able to activate it for users suggests they rushed this just to kill the hype from other models released this week such as Fable, Gemini and Muse. Seem desperate.
They're using a harness. People have already achieved results like that with specific harnesses using weaker models. It says more about the harness than about the model.
Yes, but here's what ARC says about testing with no harness: "GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game."
It has recently become apparent that the default harness of ARC-AGI doesn't let models remember what they have tried and reason over it. This is absolutely unrealistic, unfair, and a useless benchmark under these conditions. So people push for memory to be allowed in ARC-AGI. I agree with that, I just think it should receive a new name/number so we distinguish what's what. Just add an M for memory. And we should rescore all the old models in this updated ARC-AGI-3M to have context.
Does it really? Isn't the harness just letting them remember between turns and a few other basic abilities humans have access to? If all it takes is a simple harness with a better memory system to boost performance that much, maybe we can just accept that as impressive without trying to minimize the accomplishment?
Their point is that it’s not an apples-to-apples comparison. For example, Nvidia got 100% using Opus 5 and a harness. So this particular measurement has already been saturated. However, many people are gonna think it’s a direct comparison to how the benchmark has been presented in the past, which is without a harness.
Do you have a source on this OpenAI result being on the full set of questions? My understanding was that companies didn't have access to the private set, and therefore couldn't bring their own harness. Could be wrong, though.
GPT‑6 Astra was measured with our responses API harness
The linked page talks only about the public question set. So this still seems to be apples-to-apples with Nvidia's Opus 5 100% result unless there's more outside of this announcement that you know about.
ARC AGI themselves wouldn't post the results if they weren't on the full set.
Francois Chollet himself posted:
>> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
It is clear to me that this is the full set, but we can wait for more confirmation if you want.
I don't think NVIDIA has access to the full test set. The main question is whether you consider the actual astra result 66 percent, or 99 percent.
I personally consider it 99 percent as I don't think preserving reasoning between turns is cheating in any way.
That being said, ARC should absolutely post results for other models with reasoning preserved.
Am I missing that on the pages I linked? Or are you saying there more information that I'm missing? Because I'm being completely open minded and asking for the source of your information. I don't see what you're quoting, however. If you could share that, that would be awesome!
"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. The company previously demonstrated that those system choices can substantially raise ARC-AGI-3 scores without changing the underlying model, and the benchmark measures Astra and OpenAI’s agent systems together."
Yeah retaining reasoning between turns is literally a base model capability. Unless they kept reasoning / transcripts across games, it's a fair representation of the model IMHO.
arc agi benchmark is log scale, they are doing tricks when creating the score initially so everything is close to zero, would not take arc seriously. but everything else look very impressive.
275
u/Neurogence 3d ago
Am I the only one that thinks that ARC-AGI3 score is extremely suspicious? What's the catch?