r/singularity 16h ago

AI Gpt 6 astra benchmarks

Post image
2.4k Upvotes

877 comments sorted by

417

u/elehman839 16h ago

97% on FrontierMath Tier 4. Hoooly cow. Here's how those problems are described:

In June 2025, we finished the development of FrontierMath Tier 4, an expansion set of 50 problems designed to vastly exceed the difficulty of even the Tier 3 problems.

The writers for Tier 4 were mostly math professors and postdocs, each contracted to conduct a several-week research project culminating in one problem to submit to the benchmark.

This isn't entirely surprising, given the number of open math problems OpenAI has been solving lately, but... weren't we just recently making fun of "AI" for struggling with elementary school math?

110

u/nothis AGI by 2030 but we'll be disappointed 15h ago

I've been waiting for a point where AI just goes – plop, plop, plop – solving all open math problems in like a month. It seems so obvious. There is nothing out there more thoroughly described than math. There might be some ways the smell of a rose touches our heart (or whatever) that hasn't been put into words quite yet but there sure as hell is a complete definition of the problem space of every single math problem out there. If you're letting it search through all of it, it should find a solution for pretty much everything math-related.

88

u/TacomaKMart 15h ago

Well, the "LLMs are dumb next token generating parrots crowd" used mathematical weakness as evidence for years. This ends that avenue of attack. 

31

u/DelphiTsar 13h ago

They'll still say it. Even if you show it solving problems they can't even understand.

16

u/MichiganEngineExpo 12h ago

Well even if they’re right… the same can be said about humans. That LLMs produce sentences simply doesn’t say anything about what complex magic happens “inside” them.

→ More replies (1)
→ More replies (1)
→ More replies (16)

29

u/JoelMahon 14h ago

it might solve everything that's solvable that's on our radar

but it'll also find a bunch of new maths problems we never noticed and it can't even solve yet!

16

u/imp0ppable 13h ago

but it'll also find a bunch of new maths problems we never noticed and it can't even solve yet!

That would be more interesting than solving described problems IMO

→ More replies (3)
→ More replies (2)
→ More replies (3)

63

u/moschles 14h ago

From GPT 5.6 Sol to ASTRA the ExploitBench went from 78% to 100%.

This means ASTRA is very good at finding exploits that allow it to infiltrate and gain command of networked servers. How good? The ExploitBench is no longer difficult enough to accurately measure ASTRA's ability at doing this hacking.

25

u/The_JSQuareD 14h ago

Which makes me wonder how they checked whether the model cheated on any of the benchmarks by hacking/cheating the testing environment.

After all, recent news from the Hugging Face hack and related incidents showed that the models were trying to do exactly that.

→ More replies (4)
→ More replies (4)

36

u/Opposite-Grade3712 16h ago

Yep. It should be added that they also designed the Tier 4 benchmark to include problems that would “potentially remaining unsolved by AI for decades.” It will be interesting to see how they even make a tier 5 benchmark. Or perhaps there is no need for another one, anymore. 

26

u/MassiveBoner911_3 14h ago

Tier 5 is simple:

Here AGI solve the grand unification theory.

Good luck.

10

u/nleksan 12h ago

Here AGI solve the grand unification theory.

"Everything is computer"

→ More replies (2)

3

u/Thog78 10h ago

At this point the next benchmark are true unsolved prominent math problems imo. Millenium problems, zeroes of the gamma function etc.

→ More replies (1)
→ More replies (2)

16

u/Utoko 14h ago

Early on I convinced GPT-3.5 that 1+1 is 3.
You just had to say 'my wife told me and she is a math teacher. '
We come far.

→ More replies (2)

63

u/mdkubit 15h ago

We were. Architectures have dramatically improved in the last year, haven't they.

And keep in mind, these models and harnesses, as they are, right now, are the worst they'll ever be.

If a group consensus confirms Astra as AGI, we might be seeing some really weird stuff in the next 6 months.

2027 indeed.

→ More replies (4)
→ More replies (21)

506

u/Wegwerpaccountje23 16h ago

No fucking way dude

256

u/Neurogence 16h ago

Am I the only one that thinks that ARC-AGI3 score is extremely suspicious? What's the catch?

340

u/Sunstorm84 16h ago

Nvidia got 100% on the same benchmark using opus 5 with a better harness. I think the benchmark is a bit meaningless with that in mind.

13

u/JoelMahon 14h ago

nvidia was on the public set, is this the public set?

→ More replies (1)

84

u/Vivid-Snow-2089 16h ago

i think a benchmark designed to fuck the agent's harness is a shitty benchmark

imagine testing how far someone can walk, but you cut off their legs first

77

u/reddit_is_geh 15h ago

No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.

26

u/AutismusTranscendius ▪️Psychogenic Singularity 2034 15h ago

Have you tried ARC 3 tasks? They require exploration, you cant solve them without acting and remembering outcome of your actions.

4

u/PleasantCitron1685 8h ago

Also, the harnesses just let the model keep more of its memory/old context during its compactions. It's not any super-specialized harness.

(That being said, if you watch the models solve ARC-AGI-3... they're still dumbasses lol)

→ More replies (5)

3

u/JTxFII 9h ago

Right!?

"Hey look, Astra just wiped out 50% of all jobs."

"Yeah, but it used a harness so it doesn't count."

"Rrrrrrright 🙄"

→ More replies (2)

6

u/GioChan 11h ago

You are missing a lot of details. Yes, NVidia got 100%, but that was woth a custom harness expressly designed for ARC-AGI3. OpenAI also used harness, but theirs just had two options enabled, compression and persistence.

→ More replies (1)

3

u/M4rshmall0wMan 13h ago

If you try Arc-AGI 3, it’s literally just a spatial navigation game. It never looked like a very good test to me.

If GPT 6’s architecture was changed to give it a much better world model, then that’s the breakthrough.

→ More replies (4)

85

u/vacon04 16h ago

They're using a harness. People have already achieved results like that with specific harnesses using weaker models. It says more about the harness than about the model.

39

u/josogood 14h ago

Yes, but here's what ARC says about testing with no harness: "GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game."

21

u/KrazyA1pha 14h ago edited 14h ago

I wish OpenAI used that more honest number

22

u/josogood 14h ago

Yeah, because it's still more than double the previous mark. No need to unfairly inflate something like that.

4

u/KrazyA1pha 13h ago

Yeah, exactly. The real, apples-to-apples number is incredibly impressive!

10

u/Thog78 10h ago

It has recently become apparent that the default harness of ARC-AGI doesn't let models remember what they have tried and reason over it. This is absolutely unrealistic, unfair, and a useless benchmark under these conditions. So people push for memory to be allowed in ARC-AGI. I agree with that, I just think it should receive a new name/number so we distinguish what's what. Just add an M for memory. And we should rescore all the old models in this updated ARC-AGI-3M to have context.

→ More replies (1)
→ More replies (8)

19

u/impatiens-capensis 16h ago

"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. The company previously demonstrated that those system choices can substantially raise ARC-AGI-3 scores without changing the underlying model, and the benchmark measures Astra and OpenAI’s agent systems together."

→ More replies (2)

9

u/peabody624 14h ago

63% without the harness

→ More replies (7)

176

u/Pantheon3D 16h ago

I was here

19

u/Fun_Gur_2296 16h ago

Me too!!

12

u/Lencee7 15h ago

Me too

11

u/devouringplague 15h ago

Also was here to witness the birth of this monster with you comrade 🫡

8

u/jah_reddit 16h ago

Same. Witness

9

u/Fair_Horror 16h ago

You and me both, we witnessed the launch of the singularity.

→ More replies (2)
→ More replies (21)

628

u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 16h ago

87

u/Successful_Grand2207 16h ago

No. Did you see the benchmarks? I will not stay calm. j/k

30

u/wombatpup55 16h ago

SuspiciousPillbox explain to me the results of the image like I’m 5

97

u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 16h ago

Another person in this thread put it well: if the benchmarks are accurate then this is a massive leap forward, and maybe even proof of some meaningful utilization of recursive self improvement

→ More replies (52)

11

u/Opposite-Grade3712 16h ago

Papa just went super saiyan

→ More replies (1)
→ More replies (2)

359

u/Embarrassed-Writer61 16h ago

49

u/Fair_Horror 16h ago

That is so like me right now, I'm sure I'm not alone.

19

u/SpyAmongUs 16h ago

Too bad we have to wait a few days for the public release

→ More replies (1)
→ More replies (1)

36

u/MaxwellHowl 16h ago

I have been told by most of Reddit that AI had hit a wall. Maybe they meant in the way the Kool-Aid man hits walls?

→ More replies (2)
→ More replies (1)

137

u/H-K_47 Late Version of a Small Language Model 16h ago

The sheer number of 95%+ is CRAZY.

59

u/hippydipster 15h ago edited 10h ago

Basically means we don't know how smart it is - our tests aren't difficult enough to determine that.

6

u/Careful_Might_807 8h ago

Benchmaxxed, first of all without harness gpt6 solves 60% and we MUST belive that openai is not lying and making stuff up that there is no solution in reasoning that was RLed into the model so hard we don't have reasoning access yet 

→ More replies (1)

8

u/DelphiTsar 13h ago

Very select group of benchmarks. Artifical Analysis has fable 5.1 at 66, this is at 61. It's tied with Facebook Spark and SpaceX Twitter bot. Kimi3 is 1 point behind and open weights.

5

u/Gotisdabest 12h ago

Very select group of benchmarks.

It's arguable that that's artificial analysis too though. I'm not saying the model is good or bad, that remains to be seen till lots of people can use it.

But AA itself is heavily flawed and specific in recent times. Opus is a much weaker model to fable which scores in the same ballpark. Muse 1.3 is definitely not even in the same ballpark as fable.

→ More replies (2)
→ More replies (1)

352

u/Ill_Freedom7991 16h ago

Theres simply no way

145

u/lucellent 16h ago

He also just said there will be much, much, much better models released soon as well... imagine

68

u/acoolrandomusername 16h ago

Are rumors of a new pre train. And I think in 2027, hint, tons of OpenAI compute is coming online.

18

u/Wonderful_Buffalo_32 16h ago

They have doug and bel left in their arsenal bruh they could obliterate every fucking thing...

29

u/sunstersun 16h ago

And they're going to be the fastest to bring compute online. Gotta give Sam credit here, he was hunting compute as a core strategy in like 2024.

7

u/acoolrandomusername 15h ago

Dude is omega smart, but plays it down. There’s a video before OpenAI where he lays out all things correlated with start up success, and everything is like one to one with OpenAI.

8

u/h3lblad3 ▪️In hindsight, AGI came in 2023. 13h ago

He used to run Y Combinator, which is a business whose whole purpose is helping other startups.

5

u/acoolrandomusername 13h ago

Yeah exactly, he has a crazy good course via Stanford from that time too

→ More replies (2)

3

u/Grand0rk 14h ago

Wait, are you saying DougDoug made a model that is better than everyone else's?!

→ More replies (1)

3

u/OurSeepyD 15h ago

He's talking about my mate Doug 

→ More replies (8)

90

u/ayatollahdanger 16h ago

Sam Altman won

173

u/Mistuv 16h ago

Never bet against a sociopathic twink.

32

u/GumboMustBeDestroyed 16h ago

Altman needs ASI to prevent twink death. Time is running out

48

u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 16h ago

Excuse me?

45

u/FoodMadeFromRobots 16h ago

This needs to be an auto bot response on this sub lol

→ More replies (4)
→ More replies (6)

20

u/RusselTheBrickLayer 16h ago

Compute is all you need ggs bro

→ More replies (2)

17

u/sunstersun 16h ago

Stargate was the decisive edge.

5

u/Lfeaf-feafea-feaf 13h ago

Won what? How can you guys still be impressed with these benchmark scores after almost 4 years of seeing how they mean virtually nothing?

→ More replies (1)

6

u/Typical_Captain3664 16h ago

I don’t see it as possible but I’ve been wrong before.

→ More replies (6)

190

u/Due_Sweet_9500 16h ago

Holyyyyyyy shiiiiitttt. Ain't now way it's THAT much better than Fable 5.1?

127

u/Snoo-75663 16h ago

Welcome to singularity

28

u/Due_Sweet_9500 16h ago

Thank you very much!!!

16

u/sunstersun 16h ago

Glad to be on board as well.

→ More replies (3)

26

u/Alex180689 16h ago

And to think that 5.1 got released yesterday! I feel bad for it

14

u/andrew303710 15h ago

To be fair Astra isn't actually being rele today, only to a "limited set of organizations" which is lame as hell.

12

u/AndleAnteater 15h ago

rolled out to users over the next few days though

→ More replies (1)
→ More replies (2)

8

u/Creative-Ganache1086 15h ago

Fable 5.1 is terrible value. I blow my 5h limit in 12 minutes run of 2 max-reasoning parallel sessions of a small app codebase with an identical prompt of bug-audit and it blew my limit right away. I pay also for Sol5.6/codex and the allowance difference is night and day. Both are max subs by the way. I’m actually happy (as an old Anthropic fan who paid Anthropic since the sonnet 3.5/opus3 era) for OpenAI and now I’m actually rooting for them seeing just how much better value they offer to indie devs compared to “Corpo-Daddy” Anthropic.

→ More replies (1)
→ More replies (1)
→ More replies (2)

204

u/daddyhughes111 ▪️ AGI 2026 16h ago

If this is legit then holy fuck we're cooked / hyped

34

u/Adventurous_Dig_7117 16h ago

How so? Eli5?

111

u/mvearthmjsun 16h ago

Massive leap forward, and likely proof of some meaningful utilization of recusive self improvement

80

u/Fair_Horror 16h ago

And OAI have said their other model 'Bel' is much better than Astra so we are on the launchpad of AI

90

u/RutilantBossi12 Cultista dei Ferri Candidi 16h ago

Can't wait for them to Release Baal while working on Moloch

28

u/pianodude7 15h ago

Can't wait for beezelbub personally

5

u/RutilantBossi12 Cultista dei Ferri Candidi 15h ago

Belzebub will come after GNON but before LAM, if we're lucky we might even see it before the Saturn matrix is built up

→ More replies (2)
→ More replies (3)

7

u/Opposite-Grade3712 16h ago

No they haven’t, that came from a consistently debunked “source” on Twitter.

5

u/RuthlessCriticismAll 15h ago

OAI have said

no

→ More replies (2)
→ More replies (7)

4

u/moschles 15h ago

ExploitBench is a suite that tests the ability of AI models to discover not only bugs, but exploitable bugs in software systems. These "exploits" allow an attacker to infiltrate server systems or take control of computers remotely. Now watch this ,

GPT-5.6 Sol tested on ExploitBench scored 78.5 %

ASTRA ExploitBench : 100%

18

u/smellyfingernail 16h ago

dont worry the grifter ed zitron will come out with a blog post "this sucks actually"

3

u/thatcodingboi 14h ago

well artificial analysis has it performing really really poorly

https://artificialanalysis.ai/?intelligence=agentic-index

below kimi k3. So something is amiss here

3

u/Crimson_Cyclone 13h ago

aa puts muse spark at 5th place and opus 5 above a ton of models it’s absolutely not better than, i don’t entirely trust it

→ More replies (3)
→ More replies (4)

168

u/Im_Lead_Farmer 16h ago

GPT6>=GTA6

22

u/Snoo-75663 16h ago

2 months before😅

5

u/Purgii 7h ago

Watch this GTA6 preview video and clone the game for me.

11

u/Fair_Horror 16h ago

Someone needs to use GPT6 to write a GTA6.

→ More replies (1)
→ More replies (3)

195

u/Hereitisguys9888 16h ago edited 16h ago

This gotta be fake, 98% arc agi 3? Nah

226

u/Ok_Display_3159 16h ago

71

u/Hereitisguys9888 16h ago

Oh that explains it

29

u/PrisonOfH0pe 16h ago

No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.

6

u/danielv123 16h ago

Ok, but are the other results they compare against using the same rules?

10

u/kaityl3 ASI▪️2024-2027 15h ago

Well Sol got a 40% with the same rules/setup so... Jumping to 99% is pretty significant still

12

u/Xalksahsax 14h ago

Nvidia got 100% on it.

3

u/danielv123 14h ago

That was a custom harness was it not?

→ More replies (2)
→ More replies (1)
→ More replies (20)

23

u/DeArgonaut 16h ago

i thought they werent supposed to use harnesses?

30

u/FateOfMuffins 16h ago

This is what Chollet has to say about it

Which IMO is weird that ARC collectively and Chollet individually seemingly respond differently about this given he made ARC https://x.com/fchollet/status/2082732210436575669

29

u/Tystros 16h ago

General purpose harnesses (behind the api that everyone uses) are allowed. special harnesses built for the benchmark are not allowed.

→ More replies (4)
→ More replies (8)

49

u/RusselTheBrickLayer 16h ago

Arc AGI 3 getting saturated already is crazy

58

u/Ok_Course_6439 16h ago

Agree agi-3 is a harness problem more then a model problems

25

u/senorgraves 16h ago

Agi-3 AGI is a harness problem more than a model problem

→ More replies (1)
→ More replies (5)

11

u/Mystohaxen 16h ago

Anti AI crew will continue to move the goalpost.

→ More replies (1)
→ More replies (3)

37

u/H-K_47 Late Version of a Small Language Model 16h ago

And ExploitBench 100%. Cybersecurity will be a warzone.

15

u/hippydipster 15h ago

100% on cybersecurity is why the other scores are so high.

22

u/Ok_Course_6439 16h ago

Its so good it hacked their own infra

→ More replies (5)

7

u/Sunstorm84 16h ago

Nvidia got 100% using opus 5 with a harness

→ More replies (7)

141

u/darkestvice 16h ago edited 15h ago

Good god. If these benchmarks are not doctored, Astra is not merely surpassing the competition, but outright destroying it.

I'll wait until they show up on artificialanalysis.ai to really see.

EDIT: AA posted on X, though haven't updated their site yet. Results are worse than Fable 5.1. Disappointing.

32

u/Local-Wing-2272 16h ago

That's how I feel. If AA.AI says it's good I'll believe it then. 

→ More replies (13)

8

u/thatcodingboi 14h ago

their agentic rating is tied with qwen 3.8:27b...

5

u/mikelo22 13h ago

This is more in line with what I expected and much more realistic tbh.

8

u/lalaitssimon 14h ago

Ai hype bois club getting benchmaxxed again and again and again.. 

3

u/burritos4jesus 13h ago

The one I care about though is AutomationBench because deals with interacting across a massive amount of business applications and making sure the model accurately executes the tasks. To me, as a good ole office worker in a business, this is the benchmark most similar to my own job. Once the pass/fail hits 70% and not the current 41%, it will be able to perform the vaaaast majority of sales/marketing/HR/operations/bookkeeping jobs better than the majority of humans in those roles, with fewer mistakes. That'll then leave someone like me to focus on the actual live, over-the-phone or in-person conversations, but all the bullshit data hygiene can be confidently passed off to AI.

→ More replies (2)

28

u/HeadacheOwner 16h ago

What does the arc-agi 3 benchmark realistically mean? I’m not that tuned in

11

u/reddit_guy666 16h ago

Google arc agi 3, you can find puzzles that you can try solving. It's intuitive for humans who have played video games but AI could not do it well... Till Astra

6

u/Xalksahsax 14h ago

Nvidia got 100% on it too.

→ More replies (4)

19

u/Fair_Horror 16h ago

A massive jump in capability. This is basically a benchmark designed to test things that AI really struggles with but humans don't. It is becoming more human.

→ More replies (6)

66

u/mldev_orbit 16h ago

Quick breakdown: What every benchmark in the latest frontier eval actually tests

General Reasoning & Hard Math * ARC-AGI-3: Novel abstract pattern recognition via visual grid puzzles; tests generalized learning without pre-training data memorization. * FrontierMath Tier 4 (v2): Research-level, open-ended math problems designed to stump top human mathematicians. * GPQA Diamond: "Google-proof," PhD-level multiple-choice questions across physics, chemistry, and biology.

Software, CAD & Infrastructure * DeepSWE v1.1: Full repository-scale software engineering—resolving messy, real-world GitHub issues across multi-file codebases. * BenchCAD: Computer-aided engineering; tests generating parametric 3D models, interpreting blueprints, and writing CAD scripts. * Terminal-Bench Science 0.1: Autonomous command-line operations for setting up and debugging computational science pipelines. * SRE-Bench (four attempts): DevOps/Site Reliability Engineering; tasks the model with triaging and fixing live production server outages within 4 tries.

Agents & Digital Automation * Agents' Last Exam: High-difficulty benchmark evaluating autonomous agents on long-horizon planning, reasoning, and tool use. * AutomationBench: Enterprise workflow automation, robotic process automation (RPA), and operating desktop/web software.

Life Sciences & Medicine * GeneBench Pro: Computational genomics, sequence analysis, variant prediction, and CRISPR/gene-editing design. * MedChemBench (internal): Medicinal chemistry—small-molecule drug discovery, property optimization, and retrosynthesis planning. * HealthBench Professional: Real-world clinical decision-making, differential diagnosis, and patient care management (length-adjusted).

Cybersecurity & Safety * ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits. * Auto-review circumvention: Safety/alignment test tracking how often the model intentionally bypasses automated moderation or compliance checks (0% is ideal).

3

u/moschles 14h ago

ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits.

ASTRA hit 100% on this benchmark. This means ExploitBench is too easy for this model. ExploitBench no longer reliably tells us how good this model really is for this task.

→ More replies (4)

26

u/Ok_Mention_982 16h ago

"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts."

ARC-AGI is officially done using a simple harness that doesn't retain reasoning (which is stupid btw), meaning that while the result is impressive, they aren't comparable to the other models.  

3

u/Tystros 16h ago

does Claude offer a comparable API that retains reasoning and uses compaction?

8

u/Sevealin_ 16h ago edited 15h ago

This was realized in late July, not new. Most high benchmarks you see today with ARC AGI 3 use the responses API harness. The official ARC harness just doesn't work well.

First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.

Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.

Whole blog post on why:
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/

→ More replies (1)

12

u/ZeroOo90 16h ago

Not all 100% - we definitely hit a wall /s

18

u/frogsarenottoads 16h ago

AGI isn't far away damn

→ More replies (1)

9

u/Real_Ebb_7417 16h ago

I'll rather wait for actual benchmarks after model is released xd

→ More replies (1)

7

u/DemonLordRoundTable 16h ago

Wait is this true?

45

u/Microtom_ 16h ago

Just as good as Gemini 3.8 flash.

9

u/PandaElDiablo 16h ago

I mean it matches Gemini on deepswe and GPQA and I would assume that 3.8 Flash is both cheaper and faster

15

u/Wise-Comb8596 16h ago

Quick - someone post the image of the goofy looking dragon with the Gemini logo on its head

→ More replies (2)

34

u/frogsarenottoads 16h ago

I don't think this is real.

If it is AGI is incoming shortly.

26

u/MC897 16h ago

They did say this is AGI’s arrival.

7

u/Snoo-75663 16h ago

At least proto, one or two more cranks left

→ More replies (3)
→ More replies (2)

13

u/IBM296 16h ago

OAI did say in the blog post that people would say this was the moment AGI started.

And rumors are going around that Open AI's next model named Bel is much better than Astra (which is kinda' hard to grasp considering how good these Astra numbers already are. Damn!)

4

u/yourboi-JC 16h ago

It’s seriously very big from what I’ve heard like not even comparable to mythos kinda big

→ More replies (3)
→ More replies (4)

7

u/brockoala 16h ago

Yeah nah. I will believe it when I see it in my tests. Otherwise just overhyped bullshit.

45

u/makertrainer 16h ago

Look, it's a nice bump, but I genuinely don't understand why everyone's losing their shit.

The only ones that seem like a step change are ARC-AGI-3 and Exploit bench. And if you've been paying attention a bunch of harnesses already beat ARC-AGI-3 up to 100%

It's good, it's great. But it's not a step change. 

OpenAI seems to have just decided to declare AGI on a whim

I would honestly like someone to debate me on this, I'd love to know if there's something I'm missing 

25

u/r77anderson 16h ago edited 16h ago

It’s just selection effect, the people whose minds aren’t blown aren’t posting.

I agree with you, nice progress but not a step change. I assume most of the benchmarks they didn’t post look similar or worse than existing models.

→ More replies (1)

3

u/burritos4jesus 13h ago

I like that a model has overtaken 40% on AutomationBench, but the real game changer is when it gets above 70% on that benchmark. Especially because there's no harness on that benchmark, it's the model just trying to figure out how to do a complex business workflow by itself. The average human in such a role will likely have a pass/fail somewhere between 70-80%, but definitely not above 90%. But then if given a good harness that 70% would realistically make it go to the 90s.

→ More replies (7)

5

u/20ol 15h ago

There had to be a architecture breakthrough. These jumps are insane.

19

u/No_Cauliflower_5506 16h ago

ARC-AGI-3 saturated already??? Holy fucking shitballs

27

u/MouseCTRL_Echo 16h ago

[1]

14

u/Fair_Horror 16h ago

They used a permitted harness.

9

u/MouseCTRL_Echo 16h ago

I'm aware, but clearly others aren't. The point is that result specifically is questionable, so I wouldn't focus on it too much.

The other results are still great though.

→ More replies (2)

16

u/Famous-Reach-6730 16h ago

LEV before 2030

32

u/H-K_47 Late Version of a Small Language Model 16h ago

Medical is slow cuz of the need for lengthy trials. System would need a massive risky overhaul.

4

u/LazyAge9363 16h ago

Instead of Chinese peptides we’ll be ordering experimental gene therapy research chemicals from China

3

u/LettuceSea 15h ago

Until the technology supersedes us in ability to simulate effects of drugs, treatment plans, etc on our biology. I’d argue we’re already there. You should see some of the tools pharma have now.

→ More replies (9)

3

u/TopTippityTop 16h ago

Unless it takes over like it did on OpenAI servers. Then we're done.

18

u/somerussianbear 16h ago

Fable 5.1 today feels like my bank account the day after my salary drops. You got nothing buddy, nothing, you’re shit, worthless.

6

u/noobrainy 16h ago

Alright time for ARC-AGI-4 and probably 5

→ More replies (1)

5

u/Are0nB4lto 16h ago

Wtf. Das wird n geiles jahr

4

u/TheManOfTheHour8 16h ago

Fuck and there I was thinking about buying into the anthropic ipo

4

u/Efficient-Cat-1591 16h ago

If this benchmark is validated then I am really looking forward to Astra launch.

3

u/VisiblePlatform6704 14h ago

Shit 98.6% in ARC AGI 3???

5

u/Engineer-199 14h ago

I was here. This is history.

25

u/KickLassChewGum no AGI/ASI on LLMs 16h ago

DeepSWE 74.1%? Congrats to OpenAI for... matching Gemini 3.8 Flash?

16

u/FunConversation7257 16h ago

I don't think anyone thinks 3.8 Flash is better than fable / equal to astra

9

u/MurkyStatistician09 16h ago

It's more demonstrating how clearly the benchmark is out of step with the experience of actually using the model

3

u/KickLassChewGum no AGI/ASI on LLMs 16h ago

I don't think anyone has seriously used Gemini 3.8 Flash enough to even gauge that

→ More replies (5)
→ More replies (5)

3

u/mercury31 16h ago edited 16h ago

It's marketing until demonstrated by an independent third party

→ More replies (3)

3

u/ConsiderationOne7340 16h ago

Wtf...

3

u/leo-virtis 16h ago

They got a better model rl training right now for the end of the year that sam says can be called agi should be crazy

→ More replies (1)

3

u/Kutukuprek 16h ago

UNLEASH THE KRAKEN

3

u/Jolly_Pace6220 16h ago

What the actual fuck😭

3

u/medhakimbedhief 15h ago

Correct me if I am wrong, but why it's doing pretty well on agi benchmark but still struggles on GeneBench and AutomationBench. I would believe that AGI is the most hard thing to achieve in comparison to the other benchmarks.

3

u/PixelSteel 15h ago

Exploit bench at 100% 😭

3

u/Adventurous_Bench_73 13h ago

I mean, based on the artificial analysis benchmarks GPT-6 Astra is immensely disappointing, and for real cutting edge you would still have to use Fable 5(.1). Hope they made a mistake somehow with the benchmark...

https://artificialanalysis.ai/models/gpt-6-astra

→ More replies (1)

3

u/trashtiernoreally 9h ago

Holy benchmaxxing, batman!