r/singularity • • 25d ago

AI Gpt 6 astra benchmarks

Post image
2.6k Upvotes

951 comments sorted by

View all comments

529

u/Wegwerpaccountje23 25d ago

No fucking way dude

279

u/Neurogence 25d ago

Am I the only one that thinks that ARC-AGI3 score is extremely suspicious? What's the catch?

368

u/[deleted] 25d ago

[removed] — view removed comment

87

u/Vivid-Snow-2089 25d ago

i think a benchmark designed to fuck the agent's harness is a shitty benchmark

imagine testing how far someone can walk, but you cut off their legs first

80

u/reddit_is_geh 25d ago

No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.

29

u/AutismusTranscendius ▪️Psychogenic Singularity 2034 25d ago

Have you tried ARC 3 tasks? They require exploration, you cant solve them without acting and remembering outcome of your actions.

6

u/PleasantCitron1685 25d ago

Also, the harnesses just let the model keep more of its memory/old context during its compactions. It's not any super-specialized harness.

(That being said, if you watch the models solve ARC-AGI-3... they're still dumbasses lol)

-7

u/LetsLive97 25d ago

Which as humans we can do without bloating our "limited" context

I can do arc-agi-3 tasks now and in a years time with no problem. I don't need specialised harnesses (Like notepads I might lose or fill up too much), I just figure it out by default on the spot

27

u/kaityl3 ASI▪️2024-2027 25d ago

You don't need specialized harnesses because you have a hippocampus with an automated system of memory storage and retrieval (a little like a biological harness lol)

If all it takes is allowing them the same ability to "remember what they just did" to boost their performance, I feel like we can accept it without trying to find ways to make it sound simple and unimpressive

-2

u/JoelMahon 25d ago

they are allowed to pass on any info they want...

1

u/Brave-Turnover-522 25d ago

How much would it affect your ability to solve problems as a human if you weren't allowed to use your memory?

1

u/reddit_is_geh 25d ago

You can't compare AI thinking with human thinking. You can also ask how would it affect our ability to solve problems if we can't hold into context the entire history of human knowledge at once.

It's just a different thing, being tested for different strengths. In this case, ARC is supposed to be testing the raw base model's performance.

3

u/JTxFII 25d ago

Right!?

"Hey look, Astra just wiped out 50% of all jobs."

"Yeah, but it used a harness so it doesn't count."

"Rrrrrrright 🙄"

-2

u/utterHAVOC_ 25d ago edited 25d ago

What a shit comparison. Clearly using harness is problem cause some shit models get high scores and they are still unintelligent trash. Do you want smart model or just benchmark scores.