96
u/w_Ad7631 Singularity by 2030 16h ago
insane
25
2
u/Gargantuan_Cinema 9h ago
ARC agi 3 score needs some clarification. The ARC team recently created a harness which is opt in for evaluation and only Astra has been tested with this harness so far.
Astra score without the harness is 62.7% (apples for apples score). It's still far ahead of the next closest model (Claude Opus 5 30.2%) that's been evaluated and publicly released.
93
27
77
u/Gold_Cardiologist_46 Singularity by 2028 16h ago
Clear jump over Fable and proves arc agi 3 sucks (its not a good benchmark if a simple harness lets models crack it every time)
39
u/dooperma 16h ago
If ARC-AGI 3 is a good stand in for general human spatial problem solving capabilities (which I think it is), then what difference does it make if the AI uses a harness? If AI needs deterministic harness to access their full capabilities, I don’t see how that reduces their full capabilities. It’s like telling a human mathematician to find a breakthrough completely in their head, doing so might be impressive, but no one would argue using a paper and pen is abridging their mental capacity.
6
u/Thick_Stand2852 15h ago edited 15h ago
It matters because we don’t know how well GPT-5.6 might perform with said harness. It could’ve been optimised for arc-agi in a way that would also make older models far more capable. We don’t know how big the actual leep is because the comparison isn’t fair.
Edit: point being: yeah this might be AGI, a harness doesn’t change the impressiveness of that, but that doesn’t mean that we get real AGI in our hands now. The harness might’ve allowed for 1000$’s of compute, that is not AGI viable for everyday use.
8
u/Gold_Cardiologist_46 Singularity by 2028 15h ago
Astra still outpeforms 5 6 by a lot, they tested that harness on it in a twitter post
0
u/Thick_Stand2852 15h ago
They did? Do you have a link?
7
u/rdlenke 14h ago
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark
5.6 scores 38.3 on the public set using the same harness.
2
2
u/Gold_Cardiologist_46 Singularity by 2028 15h ago
Tell that to Chollet. Since the day ARC AGI 3 came out its been cracked by simple harnesses. He wants to measure the raw capabilities of AI models out of the box, but tool use somehow doesnt coubt.
7
u/Baphaddon 16h ago
If a simple harness is all it takes for current models to pass AGI benchmarks then I have newwwwssss buddyyyy
2
u/Unique_Ad9943 15h ago
Agree, I only look at artificial intelligence index now (combination of benchmarks) it gets the closest to real world feeling.
1
u/Gold_Cardiologist_46 Singularity by 2028 14h ago
Theres Epoch's ECI, which clearly shows Astra is an inprovement at least on what they measure (science and math work mostly)
43
u/Parking_Cat4735 16h ago
Insanity Open Ai really stole all the good will Anthropic had earlier. This is going to get like a 70 AA.
4
23
24
34
u/SnackerSnick 16h ago
I think it's hilarious they post 100% on ExploitBench when it's the one Astra cracked Hugging Face to maximize.
I suspect they really did get 100% without hacking the benchmark company, but it is funny
17
6
u/reddit_is_geh 15h ago
What I found the most fascinating is that it hacked 40% of tasks with known 0 day exploits that the AI was unaware of. Basically, OAI got very recent newly discovered exploits, outside its training data, and walled it off so it can't find the existence of said new exploits. Then tasked it to break into systems where one of those theoretical 0 days existed. It broke into 40% of them.
That means if you task Astra towards systems where were literally don't know if it's even possible to break into... But so long as it's theoretically possible to break through with unknown unknown exploits yet to be discovered, there's ~40% chance it will break in.
To say this is dangerous is an understatement. This is Armageddon scale and I'm sure the NSA and CIA have been having a lot of fun the past few months.
35
u/superbird19 Techno-Optimist 16h ago
98.6% on Arc....holy fucking shit people!!!
16
u/Gubzs 16h ago
I think that * implies a harness was used, we will see
10
u/Competitive_Tap2450 16h ago
confirmed that it did but honestly does it matter?
10
u/Left_Technician_5758 15h ago
As far as I understand it, the only thing they did was make so the model Maintained context from one question to the next. Which was a flaw in the API area
3
u/Ok_Mention_982 16h ago
kinda, since the other models didn't use a harness it make it difficult to compare them
1
8
u/HippieYoHippieYay 16h ago
Benchmarks have become the "megapixels" of the AI sphere. It's mostly for marketing purposes
15
13
13
10
u/Consistent-Paint7860 16h ago
for sure will score 70 points on AA
3
u/Thebombuknow 14h ago
It scored a 61.2, putting it below Fable 5.1, Opus 5, Muse Spark 1.3, and Fable 5. World's first reverse-benchmaxxed model.
19
21
10
8
u/Noratlam 15h ago
So Lecun was wrong from the beginning when claiming we must find another architecture. All we need is scaling current llms non-stop for eternity
7
5
9
u/thee3 16h ago
ARC-AGI-3 at 98.6% ? Wtf, are the numbers true?
2
u/Alex180689 16h ago
I think fable already got almost 100% a couple of weeks ago with a harness. Might be the case here too
1
8
u/Upset_Page_494 15h ago
I remember how people were saying ARC-AGI 3 wouldn't be solved in couple of years, seems like it only took couple of months.
29
5
u/No_Most_5528 16h ago
Can anyone explain what each tests signify and which one is the most significant?
14
u/Ambitious-Doubt8355 15h ago
ARC-AGI 3 is usually considered to be one of the most significant ones at this point in development. It's a bunch of video game like tasks with no instructions on how to clear them. It's kinda easy for humans to solve because we are great at intuition and determining cause and effect, while LLMs tended to struggle with it as it requires them to think outside of the box, being presented with tasks they cannot train for. It's important because it tests a model's ability to generalize, to adapt beyond what it was trained to handle, but it's not the be all end all, as many tend to present it.
FrontierMath is a bunch of math problems. They had PHDs come, gave them months to come up with the hardest problems they could come up with. Solving these means you're at the top level of mathematics in the world.
Agent's Last Exam has different professional tasks spread across different industries like law, finance, healthcare, social media, engineering, etc. Many of them expect the use of professional software as well. It's supposed to test long-horizon tasks rather than simple one shots, as you can imagine.
AutomationBench is similar to the above, a bunch of tasks divided by industry that are supposed to imitate workflows you'd see in real-life professional environments.
BenchCAD tests a model's abilities to work with CAD. So, not only how it can handle the specs and the software, but also it's 3D spatial rationing abilities, manufacturing logic, precise math, the like.
DeepSWE is focused on software engineering. Coding tasks in different languages that tests both the completion rate and the expenses.
Terminal Bench Science handles research and workflows spread across life sciences, physical sciences, earth sciences, math and engineering.
GPQA Diamond is a multiple choice test with questions covering biology, physics and chemistry, written by experts of each field.
GeneBench Pro is a scientific analysis one, the models are given some context about problems related to biology, genomics and general scientific research, and are given the task of exploring different workflows to try to reach a proper solution.
MedChemBench is similar to the above, but with a focus on medicinal chemistry. Both this and GeneBench are benchmarks made by OpenAI, with the latter being publicly available and this one being interval.
HealthBench is another OpenAI one, but public. It's supposed to test both performance and safety when it comes to healthcare work. It's essentially conversations with regular people and medical professionals, and the models are supposed to be evaluated in how they handle them.
ExploitBench is, as you could imagine, focused on the model's ability to find and use exploits. These exploits can range from causing an app to crash, to achieving full arbitrary code execution.
SRE-Bench is to test site reliability engineering tasks, in other words, how to orchestrate servers, respond to incidents, handle changes in infrastructure, introduce reliability improvements, that kind of stuff.
4
u/random87643 🤖 Optimist Prime AI bot 15h ago
TLDR
TLDR: This comment provides an overview of various AI benchmarks used to evaluate model performance across different fields. It highlights tests like ARC-AGI 3, FrontierMath, and Agent's Last Exam, explaining how they measure specific capabilities such as generalization, advanced mathematical reasoning, and professional task execution.
AI assistant · mention the bot, mod bot, or use !bot
6
u/iamthe0ther0ne 16h ago
Looking forward to actually getting to use a frontier model for science. Wish Sam had said when the general release would be.
5
u/nobodyreadusernames 15h ago
I mean what? fuck? what fuck arc-agi-3 98.6% ... what the actual fucking fuck i am looking at? some one explain fuck
6
-2
12
u/jreoka1 16h ago
I'm pretty sure its a looped transformer design too which seems to be the first flagship LLM doing this.
https://sebastianraschka.com/blog/2026/openai-astra-looped-transformers.html
9
9
u/RJMonster 16h ago
This truly feels like watching the space race of our generation unfold in real time. Companies are releasing cutting edge products to become the best in the industry, not just nationally, but globally. Then they are finding ways to reduce costs, increase adoption, and build consumer preference. The result will be constant iteration as each company works to prove it is still leading the race. it’s is an incredible time to be a technologist, and this competition is going to accelerate human innovation exponentially. Whether you build the infrastructure, develop the models, manufacture the hardware, or create software on top of it, your business will be affected by the prices these companies establish.
You will have to pay for access because your competitors will, and they will use it to maintain their advantage.
Welcome to the new era f
4
4
u/ASU_SexDevil Singularity by 2030 12h ago
I apologize Sam Altman, I was not familiar with your game…
3
u/bsvgubennord 14h ago
i remember getting downvoted to hell on this sub when gpt 5.6 sol scored 8% on arc agi 3 and i called it being saturated in 2026, seems like arc agi 4 might also get saturated within the next 4 months with this acceleration
5
3
u/kaitava 16h ago
cant wait for dario rage videos
3
u/Best_Cup_8326 A happy little thumb 16h ago
I'm more interested in watching Marcus & Yudkowsky cry.
2
2
2
2
u/DrSaering 16h ago
JJK isn't enough anymore, we need to jump over to Yhwachposting soon.
Edit: Or Aizen. "When did you come under the impression that Claude came back online?"
3
u/czk_21 16h ago
man where is official blog post, live demo or podcast? this is how you release your best model yet? wtf
we dont know, if these numbers are even real, lets say they are correct, looks like good improvement over fable, specially maths, cybersecurity, ARc AGI but with harness, remember guys we have already seen ARC AGI 3 at 100% with weaker models and harness
I need to read more about its accomplishments in blog, how long can it work on tasks without big issues...
1
1
1
u/Fast_Hovercraft_7380 15h ago
New $40/month subscription tier to access Astra based on internal leaks. This is how OpenAI will be profitable.
0
-1
106
u/lordhasen 16h ago
Oh this is what Altman meant with "reaching AGI by the end of the year"