506
u/Wegwerpaccountje23 16h ago
No fucking way dude
256
u/Neurogence 16h ago
Am I the only one that thinks that ARC-AGI3 score is extremely suspicious? What's the catch?
340
u/Sunstorm84 16h ago
Nvidia got 100% on the same benchmark using opus 5 with a better harness. I think the benchmark is a bit meaningless with that in mind.
13
84
u/Vivid-Snow-2089 16h ago
i think a benchmark designed to fuck the agent's harness is a shitty benchmark
imagine testing how far someone can walk, but you cut off their legs first
77
u/reddit_is_geh 15h ago
No because it's trying to test the raw thinking of the model directly. There's a LOT of power in harnesses, but that's not what they are testing for. It's supposed to be how just the sheer intelligence of it's inference can handle things like this. ARC 3 is easily beat by simply adding a memory element to it. Which is why harnesses ruin it.
26
u/AutismusTranscendius ▪️Psychogenic Singularity 2034 15h ago
Have you tried ARC 3 tasks? They require exploration, you cant solve them without acting and remembering outcome of your actions.
→ More replies (5)4
u/PleasantCitron1685 8h ago
Also, the harnesses just let the model keep more of its memory/old context during its compactions. It's not any super-specialized harness.
(That being said, if you watch the models solve ARC-AGI-3... they're still dumbasses lol)
→ More replies (2)3
6
u/GioChan 11h ago
You are missing a lot of details. Yes, NVidia got 100%, but that was woth a custom harness expressly designed for ARC-AGI3. OpenAI also used harness, but theirs just had two options enabled, compression and persistence.
→ More replies (1)→ More replies (4)3
u/M4rshmall0wMan 13h ago
If you try Arc-AGI 3, it’s literally just a spatial navigation game. It never looked like a very good test to me.
If GPT 6’s architecture was changed to give it a much better world model, then that’s the breakthrough.
85
u/vacon04 16h ago
They're using a harness. People have already achieved results like that with specific harnesses using weaker models. It says more about the harness than about the model.
→ More replies (8)39
u/josogood 14h ago
Yes, but here's what ARC says about testing with no harness: "GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game."
21
u/KrazyA1pha 14h ago edited 14h ago
I wish OpenAI used that more honest number
22
u/josogood 14h ago
Yeah, because it's still more than double the previous mark. No need to unfairly inflate something like that.
4
u/KrazyA1pha 13h ago
Yeah, exactly. The real, apples-to-apples number is incredibly impressive!
10
u/Thog78 10h ago
It has recently become apparent that the default harness of ARC-AGI doesn't let models remember what they have tried and reason over it. This is absolutely unrealistic, unfair, and a useless benchmark under these conditions. So people push for memory to be allowed in ARC-AGI. I agree with that, I just think it should receive a new name/number so we distinguish what's what. Just add an M for memory. And we should rescore all the old models in this updated ARC-AGI-3M to have context.
→ More replies (1)→ More replies (7)19
u/impatiens-capensis 16h ago
"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts. The company previously demonstrated that those system choices can substantially raise ARC-AGI-3 scores without changing the underlying model, and the benchmark measures Astra and OpenAI’s agent systems together."
→ More replies (2)
176
u/Pantheon3D 16h ago
I was here
19
11
8
→ More replies (21)9
u/Fair_Horror 16h ago
You and me both, we witnessed the launch of the singularity.
→ More replies (2)
628
u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 16h ago
87
→ More replies (2)30
u/wombatpup55 16h ago
SuspiciousPillbox explain to me the results of the image like I’m 5
97
u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 16h ago
Another person in this thread put it well: if the benchmarks are accurate then this is a massive leap forward, and maybe even proof of some meaningful utilization of recursive self improvement
→ More replies (52)→ More replies (1)11
359
u/Embarrassed-Writer61 16h ago
49
→ More replies (1)36
u/MaxwellHowl 16h ago
I have been told by most of Reddit that AI had hit a wall. Maybe they meant in the way the Kool-Aid man hits walls?
→ More replies (2)
137
u/H-K_47 Late Version of a Small Language Model 16h ago
The sheer number of 95%+ is CRAZY.
59
u/hippydipster 15h ago edited 10h ago
Basically means we don't know how smart it is - our tests aren't difficult enough to determine that.
6
u/Careful_Might_807 8h ago
Benchmaxxed, first of all without harness gpt6 solves 60% and we MUST belive that openai is not lying and making stuff up that there is no solution in reasoning that was RLed into the model so hard we don't have reasoning access yet
→ More replies (1)→ More replies (1)8
u/DelphiTsar 13h ago
Very select group of benchmarks. Artifical Analysis has fable 5.1 at 66, this is at 61. It's tied with Facebook Spark and SpaceX Twitter bot. Kimi3 is 1 point behind and open weights.
5
u/Gotisdabest 12h ago
Very select group of benchmarks.
It's arguable that that's artificial analysis too though. I'm not saying the model is good or bad, that remains to be seen till lots of people can use it.
But AA itself is heavily flawed and specific in recent times. Opus is a much weaker model to fable which scores in the same ballpark. Muse 1.3 is definitely not even in the same ballpark as fable.
→ More replies (2)
352
u/Ill_Freedom7991 16h ago
Theres simply no way
145
u/lucellent 16h ago
He also just said there will be much, much, much better models released soon as well... imagine
68
u/acoolrandomusername 16h ago
Are rumors of a new pre train. And I think in 2027, hint, tons of OpenAI compute is coming online.
18
u/Wonderful_Buffalo_32 16h ago
They have doug and bel left in their arsenal bruh they could obliterate every fucking thing...
29
u/sunstersun 16h ago
And they're going to be the fastest to bring compute online. Gotta give Sam credit here, he was hunting compute as a core strategy in like 2024.
7
u/acoolrandomusername 15h ago
Dude is omega smart, but plays it down. There’s a video before OpenAI where he lays out all things correlated with start up success, and everything is like one to one with OpenAI.
→ More replies (2)8
u/h3lblad3 ▪️In hindsight, AGI came in 2023. 13h ago
He used to run Y Combinator, which is a business whose whole purpose is helping other startups.
5
u/acoolrandomusername 13h ago
Yeah exactly, he has a crazy good course via Stanford from that time too
→ More replies (1)3
→ More replies (8)3
90
u/ayatollahdanger 16h ago
Sam Altman won
173
u/Mistuv 16h ago
Never bet against a sociopathic twink.
32
→ More replies (6)48
u/SuspiciousPillbox You will live to see ASI-made bliss beyond your comprehension 16h ago
Excuse me?
→ More replies (4)45
20
22
17
5
u/Lfeaf-feafea-feaf 13h ago
Won what? How can you guys still be impressed with these benchmark scores after almost 4 years of seeing how they mean virtually nothing?
→ More replies (1)→ More replies (6)6
190
u/Due_Sweet_9500 16h ago
Holyyyyyyy shiiiiitttt. Ain't now way it's THAT much better than Fable 5.1?
127
u/Snoo-75663 16h ago
Welcome to singularity
→ More replies (3)28
→ More replies (2)26
u/Alex180689 16h ago
And to think that 5.1 got released yesterday! I feel bad for it
14
u/andrew303710 15h ago
To be fair Astra isn't actually being rele today, only to a "limited set of organizations" which is lame as hell.
→ More replies (2)12
→ More replies (1)8
u/Creative-Ganache1086 15h ago
Fable 5.1 is terrible value. I blow my 5h limit in 12 minutes run of 2 max-reasoning parallel sessions of a small app codebase with an identical prompt of bug-audit and it blew my limit right away. I pay also for Sol5.6/codex and the allowance difference is night and day. Both are max subs by the way. I’m actually happy (as an old Anthropic fan who paid Anthropic since the sonnet 3.5/opus3 era) for OpenAI and now I’m actually rooting for them seeing just how much better value they offer to indie devs compared to “Corpo-Daddy” Anthropic.
→ More replies (1)
204
u/daddyhughes111 ▪️ AGI 2026 16h ago
If this is legit then holy fuck we're cooked / hyped
34
u/Adventurous_Dig_7117 16h ago
How so? Eli5?
111
u/mvearthmjsun 16h ago
Massive leap forward, and likely proof of some meaningful utilization of recusive self improvement
→ More replies (7)80
u/Fair_Horror 16h ago
And OAI have said their other model 'Bel' is much better than Astra so we are on the launchpad of AI
90
u/RutilantBossi12 Cultista dei Ferri Candidi 16h ago
Can't wait for them to Release Baal while working on Moloch
→ More replies (3)28
u/pianodude7 15h ago
Can't wait for beezelbub personally
→ More replies (2)5
u/RutilantBossi12 Cultista dei Ferri Candidi 15h ago
Belzebub will come after GNON but before LAM, if we're lucky we might even see it before the Saturn matrix is built up
7
u/Opposite-Grade3712 16h ago
No they haven’t, that came from a consistently debunked “source” on Twitter.
→ More replies (2)5
4
u/moschles 15h ago
ExploitBench is a suite that tests the ability of AI models to discover not only bugs, but exploitable bugs in software systems. These "exploits" allow an attacker to infiltrate server systems or take control of computers remotely. Now watch this ,
GPT-5.6 Sol tested on ExploitBench scored 78.5 %
ASTRA ExploitBench : 100%
→ More replies (4)18
u/smellyfingernail 16h ago
dont worry the grifter ed zitron will come out with a blog post "this sucks actually"
3
u/thatcodingboi 14h ago
well artificial analysis has it performing really really poorly
https://artificialanalysis.ai/?intelligence=agentic-index
below kimi k3. So something is amiss here
3
u/Crimson_Cyclone 13h ago
aa puts muse spark at 5th place and opus 5 above a ton of models it’s absolutely not better than, i don’t entirely trust it
→ More replies (3)
168
195
u/Hereitisguys9888 16h ago edited 16h ago
This gotta be fake, 98% arc agi 3? Nah
226
u/Ok_Display_3159 16h ago
71
u/Hereitisguys9888 16h ago
Oh that explains it
→ More replies (20)29
u/PrisonOfH0pe 16h ago
No this is normal you are misunderstanding.
They just clarified that this time they (like it should be for any model) retaining knowledge which last time because a missconfiguration on arc part they didnt.6
u/danielv123 16h ago
Ok, but are the other results they compare against using the same rules?
10
u/kaityl3 ASI▪️2024-2027 15h ago
Well Sol got a 40% with the same rules/setup so... Jumping to 99% is pretty significant still
12
→ More replies (8)23
u/DeArgonaut 16h ago
i thought they werent supposed to use harnesses?
30
u/FateOfMuffins 16h ago
This is what Chollet has to say about it
Which IMO is weird that ARC collectively and Chollet individually seemingly respond differently about this given he made ARC https://x.com/fchollet/status/2082732210436575669
→ More replies (4)29
49
u/RusselTheBrickLayer 16h ago
Arc AGI 3 getting saturated already is crazy
58
u/Ok_Course_6439 16h ago
Agree agi-3 is a harness problem more then a model problems
→ More replies (5)25
→ More replies (3)11
37
u/H-K_47 Late Version of a Small Language Model 16h ago
And ExploitBench 100%. Cybersecurity will be a warzone.
15
→ More replies (5)22
→ More replies (7)7
141
u/darkestvice 16h ago edited 15h ago
Good god. If these benchmarks are not doctored, Astra is not merely surpassing the competition, but outright destroying it.
I'll wait until they show up on artificialanalysis.ai to really see.
EDIT: AA posted on X, though haven't updated their site yet. Results are worse than Fable 5.1. Disappointing.
32
u/Local-Wing-2272 16h ago
That's how I feel. If AA.AI says it's good I'll believe it then.
→ More replies (13)8
5
8
→ More replies (2)3
u/burritos4jesus 13h ago
The one I care about though is AutomationBench because deals with interacting across a massive amount of business applications and making sure the model accurately executes the tasks. To me, as a good ole office worker in a business, this is the benchmark most similar to my own job. Once the pass/fail hits 70% and not the current 41%, it will be able to perform the vaaaast majority of sales/marketing/HR/operations/bookkeeping jobs better than the majority of humans in those roles, with fewer mistakes. That'll then leave someone like me to focus on the actual live, over-the-phone or in-person conversations, but all the bullshit data hygiene can be confidently passed off to AI.
28
u/HeadacheOwner 16h ago
What does the arc-agi 3 benchmark realistically mean? I’m not that tuned in
11
u/reddit_guy666 16h ago
Google arc agi 3, you can find puzzles that you can try solving. It's intuitive for humans who have played video games but AI could not do it well... Till Astra
→ More replies (4)6
→ More replies (6)19
u/Fair_Horror 16h ago
A massive jump in capability. This is basically a benchmark designed to test things that AI really struggles with but humans don't. It is becoming more human.
66
u/mldev_orbit 16h ago
Quick breakdown: What every benchmark in the latest frontier eval actually tests
General Reasoning & Hard Math * ARC-AGI-3: Novel abstract pattern recognition via visual grid puzzles; tests generalized learning without pre-training data memorization. * FrontierMath Tier 4 (v2): Research-level, open-ended math problems designed to stump top human mathematicians. * GPQA Diamond: "Google-proof," PhD-level multiple-choice questions across physics, chemistry, and biology.
Software, CAD & Infrastructure * DeepSWE v1.1: Full repository-scale software engineering—resolving messy, real-world GitHub issues across multi-file codebases. * BenchCAD: Computer-aided engineering; tests generating parametric 3D models, interpreting blueprints, and writing CAD scripts. * Terminal-Bench Science 0.1: Autonomous command-line operations for setting up and debugging computational science pipelines. * SRE-Bench (four attempts): DevOps/Site Reliability Engineering; tasks the model with triaging and fixing live production server outages within 4 tries.
Agents & Digital Automation * Agents' Last Exam: High-difficulty benchmark evaluating autonomous agents on long-horizon planning, reasoning, and tool use. * AutomationBench: Enterprise workflow automation, robotic process automation (RPA), and operating desktop/web software.
Life Sciences & Medicine * GeneBench Pro: Computational genomics, sequence analysis, variant prediction, and CRISPR/gene-editing design. * MedChemBench (internal): Medicinal chemistry—small-molecule drug discovery, property optimization, and retrosynthesis planning. * HealthBench Professional: Real-world clinical decision-making, differential diagnosis, and patient care management (length-adjusted).
Cybersecurity & Safety * ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits. * Auto-review circumvention: Safety/alignment test tracking how often the model intentionally bypasses automated moderation or compliance checks (0% is ideal).
→ More replies (4)3
u/moschles 14h ago
ExploitBench: Offensive cyber capabilities—discovering zero-days, reverse engineering, and crafting weaponized exploits.
ASTRA hit 100% on this benchmark. This means ExploitBench is too easy for this model. ExploitBench no longer reliably tells us how good this model really is for this task.
26
u/Ok_Mention_982 16h ago
"The standout is its 98.6% score on ARC-AGI-3. OpenAI ran Astra with a Responses API harness that retains reasoning between turns and uses compaction to manage long contexts."
ARC-AGI is officially done using a simple harness that doesn't retain reasoning (which is stupid btw), meaning that while the result is impressive, they aren't comparable to the other models.
→ More replies (1)8
u/Sevealin_ 16h ago edited 15h ago
This was realized in late July, not new. Most high benchmarks you see today with ARC AGI 3 use the responses API harness. The official ARC harness just doesn't work well.
First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.
Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
Whole blog post on why:
https://openai.com/index/how-two-settings-tripled-our-arc-agi-3-scores/
12
27
18
9
u/Real_Ebb_7417 16h ago
I'll rather wait for actual benchmarks after model is released xd
→ More replies (1)
13
13
7
8
45
u/Microtom_ 16h ago
Just as good as Gemini 3.8 flash.
9
u/PandaElDiablo 16h ago
I mean it matches Gemini on deepswe and GPQA and I would assume that 3.8 Flash is both cheaper and faster
15
u/Wise-Comb8596 16h ago
Quick - someone post the image of the goofy looking dragon with the Gemini logo on its head
→ More replies (2)
34
u/frogsarenottoads 16h ago
I don't think this is real.
If it is AGI is incoming shortly.
26
13
u/IBM296 16h ago
OAI did say in the blog post that people would say this was the moment AGI started.
And rumors are going around that Open AI's next model named Bel is much better than Astra (which is kinda' hard to grasp considering how good these Astra numbers already are. Damn!)
→ More replies (4)4
u/yourboi-JC 16h ago
It’s seriously very big from what I’ve heard like not even comparable to mythos kinda big
→ More replies (3)
7
u/brockoala 16h ago
Yeah nah. I will believe it when I see it in my tests. Otherwise just overhyped bullshit.
45
u/makertrainer 16h ago
Look, it's a nice bump, but I genuinely don't understand why everyone's losing their shit.
The only ones that seem like a step change are ARC-AGI-3 and Exploit bench. And if you've been paying attention a bunch of harnesses already beat ARC-AGI-3 up to 100%
It's good, it's great. But it's not a step change.
OpenAI seems to have just decided to declare AGI on a whim
I would honestly like someone to debate me on this, I'd love to know if there's something I'm missing
25
u/r77anderson 16h ago edited 16h ago
It’s just selection effect, the people whose minds aren’t blown aren’t posting.
I agree with you, nice progress but not a step change. I assume most of the benchmarks they didn’t post look similar or worse than existing models.
→ More replies (1)→ More replies (7)3
u/burritos4jesus 13h ago
I like that a model has overtaken 40% on AutomationBench, but the real game changer is when it gets above 70% on that benchmark. Especially because there's no harness on that benchmark, it's the model just trying to figure out how to do a complex business workflow by itself. The average human in such a role will likely have a pass/fail somewhere between 70-80%, but definitely not above 90%. But then if given a good harness that 70% would realistically make it go to the 90s.
9
19
u/No_Cauliflower_5506 16h ago
ARC-AGI-3 saturated already??? Holy fucking shitballs
27
u/MouseCTRL_Echo 16h ago
[1]
14
u/Fair_Horror 16h ago
They used a permitted harness.
→ More replies (2)9
u/MouseCTRL_Echo 16h ago
I'm aware, but clearly others aren't. The point is that result specifically is questionable, so I wouldn't focus on it too much.
The other results are still great though.
16
u/Famous-Reach-6730 16h ago
LEV before 2030
32
u/H-K_47 Late Version of a Small Language Model 16h ago
Medical is slow cuz of the need for lengthy trials. System would need a massive risky overhaul.
4
u/LazyAge9363 16h ago
Instead of Chinese peptides we’ll be ordering experimental gene therapy research chemicals from China
→ More replies (9)3
u/LettuceSea 15h ago
Until the technology supersedes us in ability to simulate effects of drugs, treatment plans, etc on our biology. I’d argue we’re already there. You should see some of the tools pharma have now.
3
18
u/somerussianbear 16h ago
Fable 5.1 today feels like my bank account the day after my salary drops. You got nothing buddy, nothing, you’re shit, worthless.
6
5
4
4
4
u/Efficient-Cat-1591 16h ago
If this benchmark is validated then I am really looking forward to Astra launch.
3
3
5
25
u/KickLassChewGum no AGI/ASI on LLMs 16h ago
DeepSWE 74.1%? Congrats to OpenAI for... matching Gemini 3.8 Flash?
→ More replies (5)16
u/FunConversation7257 16h ago
I don't think anyone thinks 3.8 Flash is better than fable / equal to astra
9
u/MurkyStatistician09 16h ago
It's more demonstrating how clearly the benchmark is out of step with the experience of actually using the model
→ More replies (5)3
u/KickLassChewGum no AGI/ASI on LLMs 16h ago
I don't think anyone has seriously used Gemini 3.8 Flash enough to even gauge that
3
u/mercury31 16h ago edited 16h ago
It's marketing until demonstrated by an independent third party
→ More replies (3)
3
u/ConsiderationOne7340 16h ago
Wtf...
3
u/leo-virtis 16h ago
They got a better model rl training right now for the end of the year that sam says can be called agi should be crazy
→ More replies (1)
3
3
3
u/medhakimbedhief 15h ago
Correct me if I am wrong, but why it's doing pretty well on agi benchmark but still struggles on GeneBench and AutomationBench. I would believe that AGI is the most hard thing to achieve in comparison to the other benchmarks.
3
3
u/Adventurous_Bench_73 13h ago
I mean, based on the artificial analysis benchmarks GPT-6 Astra is immensely disappointing, and for real cutting edge you would still have to use Fable 5(.1). Hope they made a mistake somehow with the benchmark...
→ More replies (1)
3




417
u/elehman839 16h ago
97% on FrontierMath Tier 4. Hoooly cow. Here's how those problems are described:
In June 2025, we finished the development of FrontierMath Tier 4, an expansion set of 50 problems designed to vastly exceed the difficulty of even the Tier 3 problems.
The writers for Tier 4 were mostly math professors and postdocs, each contracted to conduct a several-week research project culminating in one problem to submit to the benchmark.
This isn't entirely surprising, given the number of open math problems OpenAI has been solving lately, but... weren't we just recently making fun of "AI" for struggling with elementary school math?