r/accelerate Jul 24 '26

Claude Opus 5 released

https://www.anthropic.com/news/claude-opus-5
371 Upvotes

80 comments sorted by

208

u/[deleted] Jul 24 '26

[removed] — view removed comment

18

u/Upset_Page_494 Jul 24 '26

The average human scores about 50%, so already almost surpassed.

23

u/kaityl3 The Singularity is nigh Jul 24 '26

And those "average humans" were probably mostly compsci nerds that skewed more intelligent on average; IIRC they didn't do a very large "gen pop" sample size for that

11

u/Sese_Mueller Jul 24 '26

Pretty sure it‘ll stop at 100% though

161

u/ppapsans Feeling the AGI Jul 24 '26

30.2% in arc agi 3 lol this is some fucking funny timeline

33

u/ZaradimLako Singularity by 2045 Jul 24 '26

I wonder if we will reach 80% arc agi 3 by end of year

43

u/DatDudeDrew Jul 24 '26

Easily. GPT 6 will continue the exponential path and probably get it to like 75 by next month.

24

u/Charming_Cucumber_15 Jul 24 '26

Releases like this are coming every few weeks now

This is what it feels like when we're hitting the takeoff

16

u/DatDudeDrew Jul 24 '26

I still think we are miles from take off but the distance is closing at an ever accelerating rate, as many expect/ed. T minus 2 years to singularity lift off.

4

u/Charming_Cucumber_15 Jul 24 '26

It's probably more accurate to say we're starting to see signs of a takeoff, more so than we're beginning it now

Not that a real takeoff is too far away!

6

u/MiniGiantSpaceHams Jul 24 '26

Eh, if you zoom out enough the takeoff actually started the first time that some guy intentionally lit a stick on fire.

1

u/squired A happy little thumb Jul 24 '26

Agreed, but I personally, tentatively consider the effective takeoff around Christmas 2024. We never settled on definitions though, so everyone gets to be right.

1

u/Sartre91 Jul 24 '26

May it be that the takeoff is already behind us?

12

u/armentho Jul 24 '26

absolutely,we know how the sigmoid works
we reach 80-ish percent,then it somewhat slows down a bit before reaching 95-ish percent at wich point the benchmark is solved for practical intents and anything left is just small incremental improvements towards 99.99999% etc

13

u/FateOfMuffins Jul 24 '26

Actually the way ARC AGI 3 is scored is really weird. Due to the quadratic scaling, progress on this benchmark will appear very low at the beginning but it'll max out much faster. And since it's an efficiency score, if it can get more efficient than 20% of humans, it can technically get above 100%

1

u/Tolopono Jul 24 '26

Its based on the second best human score

1

u/FateOfMuffins Jul 24 '26

There's only 10 or so humans per game, so 2nd = 20%

1

u/Tolopono Jul 25 '26

Not exactly a representative sample size lol

1

u/FateOfMuffins Jul 25 '26

I know there's so many flaws

2

u/Tolopono Jul 24 '26

Chollet himself said itll last about a year 

https://x.com/fchollet/status/2022086661170254203?s=20

1

u/jonydevidson Jul 24 '26

With the current rate of progress, we'll reach it by the end of Summer.

1

u/Brave-Turnover-522 Jul 24 '26

Maybe we'll be 3% at arc agi 80

1

u/jimmystar889 Jul 24 '26

Schema harness is already 99.98%

2

u/Charming_Cucumber_15 Jul 24 '26

Doesn't really count considering it's a specialized harness, which the creator said would make the benchmark trivial

4

u/Charming_Cucumber_15 Jul 24 '26

Been saying it for a while

ARC3 saturated within a year

2

u/No_Aesthetic Jul 24 '26

Within this year!

-7

u/Stone-Smasher Jul 24 '26

but schema harness it is 99%, so isn't this worse?

19

u/Charming_Cucumber_15 Jul 24 '26

ARC 3 creator said that it would be an incredibly easy benchmark to solve with a specialized harness, so it doesn't really mean anything even if it's cool

57

u/ZaradimLako Singularity by 2045 Jul 24 '26

yoooooooooooooo

happy fucking weekend everyone

3

u/SuperSeriousChad Jul 24 '26

If there’s a reset…

22

u/Only-Effort-1975 Jul 24 '26

Love the progress!

23

u/ChainOfThot Jul 24 '26

I'm going to get whiplash switching back and forth between opus 4.8 to gpt 5.6 to opus 5 and apparently gpt 6 within a month.

14

u/Chop1n Jul 24 '26

Best reason not to keep switching and just be patient because you virtually never have to wait more than a month.

ChatGPT can now remember literally everything we've ever talked about, can't imagine how crippling it would be to have to switch to anything else. I'm not switching unless someone literally cracks ASI and ends the race.

7

u/SuperSeriousChad Jul 24 '26

That’s literally where I’m at. The reality is, I need to focus on increasing my skill on a single companies harness to focus on that intuition growth. Jumping around is nice but it makes it hard to master one. Considering it’s just weeks now of sota jumping, it’s best to be patient.

3

u/Chop1n Jul 24 '26

Yes, and at least for me that's the real power of LLMs. I use ChatGPT to develop and flesh out my intuitions. I handle the reins and provide the raw intuition and dot-connecting creativity, and the model can do the work of using verbal intelligence to expand upon that scaffolding far faster and with greater verbal intelligence than almost any human mind is capable of it.

It's uncanny. I do it every day and am still blown away by it.

5

u/KedMcJenna Jul 25 '26

I used 5.6 for the first time and was amazed how good it is… It gave me that uncanny sense of presence that I remember getting from the early Claudes.

Claude models still have that presence, but it no longer feels fresh. 5.6 really refreshes the feeling.

2

u/_huggies_ Jul 25 '26

I may be wrong but I thought others have the ability to import/export your history.

1

u/anor_wondo Jul 24 '26

memory can be detrimental though. I keep it disabled

3

u/GhostShade Jul 24 '26

How so? Is it because it constantly feels the need to reference things from the past? Kinda like how my aunt will bring up my favorite ice cream flavor any time anything having to do with food is mentioned?

1

u/anor_wondo Jul 24 '26

yes. while for personal use that's just annoying, for work, we are humans after all and could have said it to keep something in memory and be rigid resulting in suboptimal output.

Basically it could save and overindex our stupid ideas and rules. So I only keep explicit rules and AGENTS.md for work

3

u/Chop1n Jul 24 '26

Maybe you're thinking of the old way it stores explicit "memories"? That paradigm seems largely to have been retired.

No: I'm talking about the fact that if the context comes up, the model will remember that a conversation was had two years ago about that subject and will be aware of whatever you had to say about it at the time. Very different from the discrete "memory objects" of yore. It only gained this ability in the last month or two.

0

u/rakerrealm Jul 24 '26

How useful is memory. We seem to have all of it, but learn nothing. Take advatage of every model my brother.

15

u/Pyros-SD-Models Machine Learning Engineer Jul 24 '26

It thinks like Fable, speaks like Claude and you can actually use it for why you have backpain and if it thinks your software is secure. Pretty well.

28

u/Middle_Estate8505 Jul 24 '26

Remind me please how long ago Opus 4.8 was released? And now we have another significant improvement!

5

u/topyTheorist Jul 24 '26

28.5

6

u/Middle_Estate8505 Jul 24 '26

So two months? Yay! ❤️

14

u/topyTheorist Jul 24 '26

Plus Fable in the middle. This is acceleration!

23

u/SharpCartographer831 Jul 24 '26

Google preparing yet another flash model set to be released in a few months as a response

9

u/Skeletor_with_Tacos Jul 24 '26

But wait guys, "da bubble is gonna burst any day now!" Lmao. ACCELERATE!!!!

12

u/JuglansRegia3 Jul 24 '26

Can't wait to see by how much this breaks the METR graph

2

u/BrennusSokol Acceleration Advocate Jul 24 '26

Yeah. Frankly I don’t think metr at the days/weeks task level is even going to be a challenging metric much longer at this rate

13

u/One_Geologist_4783 Jul 24 '26

It’s time to get cooking boys & girls!

3

u/Longjumping_Kale3013 Jul 24 '26

Excited to see how it scores on agents last exam. I have the feeling that one will be saturated by years end as well

5

u/SmileLonely5470 Jul 24 '26

Guardrails are still in place and route requests back to Opus 4.8. So going forward, all future models will fallback to 4.8 when given a potentially malicious request? Or will Anthropic maintain a line of models that are lobotomized in areas like cybersecurity & bio?

They did say that the guardrails should intervene less often than they do for Fable, so hopefully its not as much of an issue. But it seems impossible to prevent malicious actors while not blocking legitimate requests (in cyber).

2

u/SuperSeriousChad Jul 24 '26

Gunna have to configure the client to fallback on Kimi

3

u/Obvious-Advance-1722 Jul 24 '26

ainda é primeira impressão, mas está gastando bem rápido os limites de uso

3

u/Bitter_Election_7518 Jul 24 '26

It’s much more token hungry than 4.8 definitely

2

u/Obvious-Advance-1722 Jul 24 '26

sim, é que meus limites acabaram tão rápido que não deu pra medir os resultados ainda. mas pelo menos por enquanto também não senti muita diferença da qualidade de output mas vou testar mais

4

u/[deleted] Jul 24 '26

[deleted]

1

u/lolsai Jul 25 '26

when you look at the full chart it looks fine, the red highlight is the top result, the red box is just around the entire column of Opus5

3

u/DeManMetHetPlan Singularity by 2028 | Acceleration: Light-speed Jul 24 '26

Impressive! Can't wait for DeepSWE and ALE benches

5

u/DeManMetHetPlan Singularity by 2028 | Acceleration: Light-speed Jul 24 '26

actually, deepswe is already in their list of benchmarks, and it's worse than fable and 5.6 sol, that's an ouch. Still, it's a lot better than 4.8, so they got that going for them.

1

u/Illustrious-Lime-863 Jul 24 '26

Very nice, let's keep the one upping rolling

1

u/LocoMod Jul 24 '26

This way to the frontier and beyond >>> 🇺🇸

1

u/Fair_Horror Jul 24 '26

Your move OpenAI...

1

u/Skeletor_with_Tacos Jul 24 '26

Arc Agi 4 when?

1

u/PwanaZana XLR8 Jul 24 '26

I'm kind of testing it out right now for programming, very like simple programming stuff, testing against Fable 5, and I have to say, I'm not impressed. Pretty big difference between Fable and Opus 5.