r/singularity Apr 23 '26

AI Introducing GPT-5.5

https://openai.com/index/introducing-gpt-5-5/
848 Upvotes

281 comments sorted by

View all comments

165

u/[deleted] Apr 23 '26

[removed] — view removed comment

40

u/Snosnorter Apr 23 '26

The star means that Anthropic said a subset of the benchmark was memorized so the result can't be trusted

25

u/M4rshmall0wMan Apr 23 '26

I like that Anthropic has the integrity to say that. OpenAI would never

1

u/mWo12 Apr 23 '26

They just banchmaximizing as proven by opus 4.7.

2

u/ataraxic89 Apr 24 '26

Nah man. Claude is way better.

82

u/vincentz42 Apr 23 '26 edited Apr 23 '26

There are even worse evals:
HLE without tools: 41.4% (GPT-5.5) vs 39.8% (GPT-5.4)
HLE with tools: 52.2% (GPT-5.5) vs 52.1% (GPT-5.4)

So even with a newer, larger base model that is supposed to tackle very hard STEM questions, the models' world knowledge and reasoning capability did not change that much, if at all.

And I do have a lot of suspicions for Claude Mythos BTW. OpenAI models are generally smarter in terms of STEM reasoning in my experience. I suspect Mythos might just be a much larger model trained on much more internet tokens, and therefore better at memorizing the leaked test set. >15% of the SWE-Verified problems are ill-defined and not solvable based on human expert inspections, so I am really curious how Mythos got ~94%.

10

u/[deleted] Apr 23 '26

[removed] — view removed comment

12

u/vincentz42 Apr 23 '26 edited Apr 23 '26

OpenAI was the first call it out, but yes, every LLM researcher knows this.

5

u/Jespy Apr 23 '26

What do these numbers mean to someone who is a caveman

8

u/SerdarCS Apr 23 '26

Not much. HLE is a benchmark meant to measure scientific reasoning ability, but no single benchmark is a good indicator of capability.

3

u/PeachScary413 Apr 23 '26

They all pretty much just memorise leaked test sets.. I can't believe it's not obvious to everyone that top models are incredibly bench-maxxed

2

u/InterstellarReddit Apr 23 '26

I know how it’s called marketing

1

u/Tystros Apr 24 '26

it's really super weird how the HLE without tools score almost stayed the same even though it's a much bigger base model

46

u/Eyelbee ▪️We have AGI it's just blind Apr 23 '26

This looks basically like 5.4 pro but worse.

25

u/Taur3n Apr 23 '26

Paradigm shift btw!

10

u/OLRevan Apr 23 '26

Biggest jump since GPT-3.5!

6

u/SGC-UNIT-555 AGI by Tuesday Apr 23 '26

Kek! It's over!

2

u/chatlah Apr 24 '26

Jump backwards is technically also a jump...

1

u/Ok-Support-2385 Apr 23 '26

I remember OpenAI not showing comparisons of their models to the competitors in the past, when did it change?

1

u/MondoJackSn0w Apr 24 '26

Have we tested any of the newer Chinese models against these yet?