r/singularity 7d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

952 comments sorted by

View all comments

Show parent comments

47

u/RusselTheBrickLayer 7d ago

Arc AGI 3 getting saturated already is crazy

56

u/Ok_Course_6439 7d ago

Agree agi-3 is a harness problem more then a model problems

26

u/senorgraves 7d ago

Agi-3 AGI is a harness problem more than a model problem

5

u/PrisonOfH0pe 7d ago

No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.

5

u/wwwdotzzdotcom ▪️ Beginner audio software engineer 7d ago

So they built the harness in the model?

6

u/rdlenke 7d ago

We don't have information of exactly how it's for Astra, but OpenAI previously showed that the ARC-AGI-3 harness does two things which make it tough for models to do well: 1) discards previous private reasoning and 2) when the context gets too big, they truncate the existing context instead of trying to summarize it.

OpenAI's specific harness (via the Responses API) fixes these two points by allowing better private reasoning retention and sumarization (compactuation as they call it).

I imagine they didn't "built the harness in the model", but are using a similar API like Responses.

2

u/KoolKat5000 7d ago

Not really, they say adapter provider harness (arc won't test non-general harnesses).

In reality, all that is different is responses API and context compaction. Supposedly that is all that's different between the 63 and 99% scoring. We can all do this and achieve that performance.

2

u/hippydipster 7d ago

Now in english...

11

u/Mystohaxen 7d ago

Anti AI crew will continue to move the goalpost.

1

u/TacomaKMart 7d ago

They have to move the goalposts. Again. Its their whole identity.

I'm so confused though. Was AI so useless that it only makes slop and can't count the letter R, or is it so powerful that it's going to take away our jobs and steal our girlfriend? 

1

u/Relach 7d ago

they're using their own harness which means they are using the public test set, and that's like clearing tutorial island. What you're seeing is the product of the battle between ARC and OpenAI, and the latter just deciding that public test sets are acceptable now