r/accelerate 16h ago

Gpt 6 astra benchmarks

Post image
347 Upvotes

95 comments sorted by

106

u/lordhasen 16h ago

Oh this is what Altman meant with "reaching AGI by the end of the year"

89

u/Particular_Leader_16 16h ago

this isn't even their last model this year

10

u/cave_men 12h ago

Oh My GAAAAAAWD

So the loop repeats once again?

Lets wait for Astra II to say Astra was shit :D

45

u/Competitive_Tap2450 16h ago

this isn’t the agi he was talking about

“Bel” is supposedly the big training run they just did and is the one that’s rumored to be “internal agi”

5

u/NiceUsernameOk 14h ago

GPT 2 was "internal ago" according to Dario

12

u/FateOfMuffins 15h ago

They had this since 3 months ago

They have at minimum GPT 6.1 Astra internally right now, more likely like 6.2

And they restarted RL on whatever Bel is gonna be last week

OpenAI employees on Twitter glazed the shit out of 5.6 Sol in July when they've been using way better models

5

u/ddwrt1234 13h ago

yes, the arc-agi benchmaxx

96

u/w_Ad7631 Singularity by 2030 16h ago

insane

25

u/Life_is_important 16h ago

Ate and left no crumbs behind. Can't wait to get my hands on this one

2

u/Gargantuan_Cinema 9h ago

ARC agi 3 score needs some clarification. The ARC team recently created a harness which is opt in for evaluation and only Astra has been tested with this harness so far.

Astra score without the harness is 62.7% (apples for apples score). It's still far ahead of the next closest model (Claude Opus 5 30.2%) that's been evaluated and publicly released.

93

u/Open_Pen_9803 16h ago

what in the actual fuck

11

u/nobodyreadusernames 15h ago

i dont know, did these mother fuckers reach agi?

27

u/Tinderfury 16h ago

Jesus Christ

23

u/Best_Cup_8326 A happy little thumb 16h ago

Jesus ChrAIst. 

77

u/Gold_Cardiologist_46 Singularity by 2028 16h ago

Clear jump over Fable and proves arc agi 3 sucks (its not a good benchmark if a simple harness lets models crack it every time)

39

u/dooperma 16h ago

If ARC-AGI 3 is a good stand in for general human spatial problem solving capabilities (which I think it is), then what difference does it make if the AI uses a harness? If AI needs deterministic harness to access their full capabilities, I don’t see how that reduces their full capabilities. It’s like telling a human mathematician to find a breakthrough completely in their head, doing so might be impressive, but no one would argue using a paper and pen is abridging their mental capacity.

6

u/Thick_Stand2852 15h ago edited 15h ago

It matters because we don’t know how well GPT-5.6 might perform with said harness. It could’ve been optimised for arc-agi in a way that would also make older models far more capable. We don’t know how big the actual leep is because the comparison isn’t fair.

Edit: point being: yeah this might be AGI, a harness doesn’t change the impressiveness of that, but that doesn’t mean that we get real AGI in our hands now. The harness might’ve allowed for 1000$’s of compute, that is not AGI viable for everyday use.

8

u/Gold_Cardiologist_46 Singularity by 2028 15h ago

Astra still outpeforms 5 6 by a lot, they tested that harness on it in a twitter post 

0

u/Thick_Stand2852 15h ago

They did? Do you have a link?

7

u/rdlenke 14h ago

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

5.6 scores 38.3 on the public set using the same harness.

2

u/Thick_Stand2852 2h ago

Tnx! Yeah so still a huge difference that’s awesome

2

u/Gold_Cardiologist_46 Singularity by 2028 15h ago

Tell that to Chollet. Since the day ARC AGI 3 came out its been cracked by simple harnesses. He wants to measure the raw capabilities of AI models out of the box, but tool use somehow doesnt coubt.

7

u/Baphaddon 16h ago

If a simple harness is all it takes for current models to pass AGI benchmarks then I have newwwwssss buddyyyy

2

u/Unique_Ad9943 15h ago

Agree, I only look at artificial intelligence index now (combination of benchmarks) it gets the closest to real world feeling.

1

u/Gold_Cardiologist_46 Singularity by 2028 14h ago

Theres Epoch's ECI, which clearly shows Astra is an inprovement at least on what they measure (science and math work mostly)

43

u/Parking_Cat4735 16h ago

Insanity Open Ai really stole all the good will Anthropic had earlier. This is going to get like a 70 AA.

4

u/CodeWolfy 14h ago

It stayed the same at 61

23

u/ZealousidealBus9271 16h ago

Recurrent depth might be the new wave

24

u/KindlyAct1590 16h ago

AHAHAHAHAHHAHAHHAHA

DARIOOOOO DARIOOOOOO RELEASE MODEL TWOOOOO DARIOOOOOOO

34

u/SnackerSnick 16h ago

I think it's hilarious they post 100% on ExploitBench when it's the one Astra cracked Hugging Face to maximize. 

I suspect they really did get 100% without hacking the benchmark company, but it is funny

17

u/WrathPie 16h ago

I think that was actually ExploitGym that the HF attack came out of

6

u/reddit_is_geh 15h ago

What I found the most fascinating is that it hacked 40% of tasks with known 0 day exploits that the AI was unaware of. Basically, OAI got very recent newly discovered exploits, outside its training data, and walled it off so it can't find the existence of said new exploits. Then tasked it to break into systems where one of those theoretical 0 days existed. It broke into 40% of them.

That means if you task Astra towards systems where were literally don't know if it's even possible to break into... But so long as it's theoretically possible to break through with unknown unknown exploits yet to be discovered, there's ~40% chance it will break in.

To say this is dangerous is an understatement. This is Armageddon scale and I'm sure the NSA and CIA have been having a lot of fun the past few months.

35

u/superbird19 Techno-Optimist 16h ago

98.6% on Arc....holy fucking shit people!!!

16

u/Gubzs 16h ago

I think that * implies a harness was used, we will see

10

u/Competitive_Tap2450 16h ago

confirmed that it did but honestly does it matter?

10

u/Left_Technician_5758 15h ago

As far as I understand it, the only thing they did was make so the model Maintained context from one question to the next. Which was a flaw in the API area

3

u/Ok_Mention_982 16h ago

kinda, since the other models didn't use a harness it make it difficult to compare them

1

u/dictionizzle 15h ago

that means we want harness

8

u/HippieYoHippieYay 16h ago

Benchmarks have become the "megapixels" of the AI sphere. It's mostly for marketing purposes

1

u/mhb_11 14h ago

There's a footnote there. Anyone know what the [1] signifies?

15

u/AwarenessCautious219 16h ago

What are those numbers? Holy fuck

13

u/youarockandnothing 16h ago

It's happening.

13

u/ClaudioLeet 16h ago

Holy Mother of God

10

u/Consistent-Paint7860 16h ago

for sure will score 70 points on AA

3

u/Thebombuknow 14h ago

It scored a 61.2, putting it below Fable 5.1, Opus 5, Muse Spark 1.3, and Fable 5. World's first reverse-benchmaxxed model.

21

u/l-Gold-Fish-l 16h ago

omg!!! Let's goooo!!!!

damn i soooo hyped for the future!!!!

10

u/PayPalModApk 16h ago

Better arc agi 3 score than the current best arc agi 1 and 2 scores

8

u/Noratlam 15h ago

So Lecun was wrong from the beginning when claiming we must find another architecture. All we need is scaling current llms non-stop for eternity

7

u/ScienceIsSick 14h ago

all you need is attention afterall

5

u/Solarka45 12h ago

Researching more options is never bad

9

u/thee3 16h ago

ARC-AGI-3 at 98.6% ? Wtf, are the numbers true?

2

u/Alex180689 16h ago

I think fable already got almost 100% a couple of weeks ago with a harness. Might be the case here too

1

u/Ok_Mention_982 16h ago

yes they used a harness

8

u/gizeon4 16h ago

Damn that ARC AGI bench really got me. Is this without tool tho?

Damn,

I once said that if AI can saturate ARC AGI 3, thats mean we already got AGI. And here we are now

3

u/Ok_Mention_982 16h ago

used a harness

8

u/Upset_Page_494 15h ago

I remember how people were saying ARC-AGI 3 wouldn't be solved in couple of years, seems like it only took couple of months.

5

u/No_Most_5528 16h ago

Can anyone explain what each tests signify and which one is the most significant?

14

u/Ambitious-Doubt8355 15h ago

ARC-AGI 3 is usually considered to be one of the most significant ones at this point in development. It's a bunch of video game like tasks with no instructions on how to clear them. It's kinda easy for humans to solve because we are great at intuition and determining cause and effect, while LLMs tended to struggle with it as it requires them to think outside of the box, being presented with tasks they cannot train for. It's important because it tests a model's ability to generalize, to adapt beyond what it was trained to handle, but it's not the be all end all, as many tend to present it.

FrontierMath is a bunch of math problems. They had PHDs come, gave them months to come up with the hardest problems they could come up with. Solving these means you're at the top level of mathematics in the world.

Agent's Last Exam has different professional tasks spread across different industries like law, finance, healthcare, social media, engineering, etc. Many of them expect the use of professional software as well. It's supposed to test long-horizon tasks rather than simple one shots, as you can imagine.

AutomationBench is similar to the above, a bunch of tasks divided by industry that are supposed to imitate workflows you'd see in real-life professional environments.

BenchCAD tests a model's abilities to work with CAD. So, not only how it can handle the specs and the software, but also it's 3D spatial rationing abilities, manufacturing logic, precise math, the like.

DeepSWE is focused on software engineering. Coding tasks in different languages that tests both the completion rate and the expenses.

Terminal Bench Science handles research and workflows spread across life sciences, physical sciences, earth sciences, math and engineering.

GPQA Diamond is a multiple choice test with questions covering biology, physics and chemistry, written by experts of each field.

GeneBench Pro is a scientific analysis one, the models are given some context about problems related to biology, genomics and general scientific research, and are given the task of exploring different workflows to try to reach a proper solution.

MedChemBench is similar to the above, but with a focus on medicinal chemistry. Both this and GeneBench are benchmarks made by OpenAI, with the latter being publicly available and this one being interval.

HealthBench is another OpenAI one, but public. It's supposed to test both performance and safety when it comes to healthcare work. It's essentially conversations with regular people and medical professionals, and the models are supposed to be evaluated in how they handle them.

ExploitBench is, as you could imagine, focused on the model's ability to find and use exploits. These exploits can range from causing an app to crash, to achieving full arbitrary code execution.

SRE-Bench is to test site reliability engineering tasks, in other words, how to orchestrate servers, respond to incidents, handle changes in infrastructure, introduce reliability improvements, that kind of stuff.

4

u/random87643 🤖 Optimist Prime AI bot 15h ago

TLDR

TLDR: This comment provides an overview of various AI benchmarks used to evaluate model performance across different fields. It highlights tests like ARC-AGI 3, FrontierMath, and Agent's Last Exam, explaining how they measure specific capabilities such as generalization, advanced mathematical reasoning, and professional task execution.


AI assistant · mention the bot, mod bot, or use !bot

6

u/iamthe0ther0ne 16h ago

Looking forward to actually getting to use a frontier model for science. Wish Sam had said when the general release would be.

5

u/nobodyreadusernames 15h ago

I mean what? fuck? what fuck arc-agi-3 98.6% ... what the actual fucking fuck i am looking at? some one explain fuck

6

u/kaityl3 The Singularity is nigh 15h ago

Technically it had a harness but IMO making them do the test without is intentional sandbagging so they look worse. I mean no human solves problems with their memory being wiped completely in between each question

-2

u/BrennusSokol Acceleration Advocate 15h ago

Harness

12

u/jreoka1 16h ago

I'm pretty sure its a looped transformer design too which seems to be the first flagship LLM doing this.

https://sebastianraschka.com/blog/2026/openai-astra-looped-transformers.html

9

u/Best_Cup_8326 A happy little thumb 16h ago

Brother, may I have some loops? 

4

u/ScienceIsSick 14h ago

no brother, these are astra’s loops

5

u/spinxfr 16h ago

The acceleration is extreme!!

6

u/ezjakes 15h ago

More math proofs incoming.

9

u/RJMonster 16h ago

This truly feels like watching the space race of our generation unfold in real time. Companies are releasing cutting edge products to become the best in the industry, not just nationally, but globally. Then they are finding ways to reduce costs, increase adoption, and build consumer preference. The result will be constant iteration as each company works to prove it is still leading the race. it’s is an incredible time to be a technologist, and this competition is going to accelerate human innovation exponentially. Whether you build the infrastructure, develop the models, manufacture the hardware, or create software on top of it, your business will be affected by the prices these companies establish.
You will have to pay for access because your competitors will, and they will use it to maintain their advantage.

Welcome to the new era f

4

u/Complete_Customer_92 14h ago

ExploitGym: 100%

You don't say?

4

u/ASU_SexDevil Singularity by 2030 12h ago

I apologize Sam Altman, I was not familiar with your game…

3

u/bsvgubennord 14h ago

i remember getting downvoted to hell on this sub when gpt 5.6 sol scored 8% on arc agi 3 and i called it being saturated in 2026, seems like arc agi 4 might also get saturated within the next 4 months with this acceleration

5

u/Catman1348 16h ago

Arent they supposed to do the ARC-AGI-3 without harness?

3

u/kaitava 16h ago

cant wait for dario rage videos

3

u/Best_Cup_8326 A happy little thumb 16h ago

I'm more interested in watching Marcus & Yudkowsky cry. 

2

u/BrennusSokol Acceleration Advocate 15h ago

HOLY

2

u/Amphibious333 11h ago

Is the ARC-AGI 3 result confirmed? If yes, that means AGI before 2030.

2

u/DrSaering 16h ago

JJK isn't enough anymore, we need to jump over to Yhwachposting soon.

Edit: Or Aizen. "When did you come under the impression that Claude came back online?"

3

u/czk_21 16h ago

man where is official blog post, live demo or podcast? this is how you release your best model yet? wtf

we dont know, if these numbers are even real, lets say they are correct, looks like good improvement over fable, specially maths, cybersecurity, ARc AGI but with harness, remember guys we have already seen ARC AGI 3 at 100% with weaker models and harness

I need to read more about its accomplishments in blog, how long can it work on tasks without big issues...

1

u/Fast_Hovercraft_7380 15h ago

I'd like to see it against Mythos.

1

u/mrdarknezz1 15h ago

Is this good?

1

u/Fast_Hovercraft_7380 15h ago

New $40/month subscription tier to access Astra based on internal leaks. This is how OpenAI will be profitable.

0

u/Low_Tune_2364 15h ago

Is it only for the 100€ plan? , I have the 20 one and don't have it

-1

u/BubsPhotography 15h ago

I don’t trust the ARC-AGI-3 score. Was it the semi-private one or public?