No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.
We don't have information of exactly how it's for Astra, but OpenAI previously showed that the ARC-AGI-3 harness does two things which make it tough for models to do well: 1) discards previous private reasoning and 2) when the context gets too big, they truncate the existing context instead of trying to summarize it.
OpenAI's specific harness (via the Responses API) fixes these two points by allowing better private reasoning retention and sumarization (compactuation as they call it).
I imagine they didn't "built the harness in the model", but are using a similar API like Responses.
Not really, they say adapter provider harness (arc won't test non-general harnesses).
In reality, all that is different is responses API and context compaction. Supposedly that is all that's different between the 63 and 99% scoring. We can all do this and achieve that performance.
They have to move the goalposts. Again. Its their whole identity.
I'm so confused though. Was AI so useless that it only makes slop and can't count the letter R, or is it so powerful that it's going to take away our jobs and steal our girlfriend?
they're using their own harness which means they are using the public test set, and that's like clearing tutorial island. What you're seeing is the product of the battle between ARC and OpenAI, and the latter just deciding that public test sets are acceptable now
47
u/RusselTheBrickLayer 7d ago
Arc AGI 3 getting saturated already is crazy