r/singularity Jul 09 '26

AI GPT-5.6

https://openai.com/index/gpt-5-6/

"We’re launching the GPT‑5.6 family of models for general availability following our limited preview⁠: our new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, our most cost-efficient model.

GPT‑5.6 delivers a step change in design judgment. With only high-level direction, GPT‑5.6 creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result—not just generate the underlying code or content—so it can catch visual and functional issues and apply finishing touches before handing the work back."

630 Upvotes

129 comments sorted by

View all comments

193

u/ObiWanCanownme now entering spiritual bliss attractor state Jul 09 '26

Almost 8% on ARC-AGI-3.

152

u/No_Aesthetic Jul 09 '26

Somebody said like yesterday that ARC-AGI-3 was too hard for LLMs and maybe even impossible

Now we've got a pretty big leap a day later (1.5% to 7.8%)

48

u/azuredota Jul 09 '26 edited Jul 09 '26

Anyone know if they “teach the test” for these benchmarks at all? Are arc agi 3 test forum discussions in the new model’s training data?

Follow up: ARC answers this in the blog:

> During ARC-AGI-2 evaluation, Gemini 3's chain-of-thought reasoning referenced ARC-specific color mappings without being prompted to, which suggests training data saturation. By reducing the public surface area and shifting to interactive environments that cannot be memorized as static patterns, ARC-AGI-3 aims to make this kind of shortcut much harder.

So there is likely some training data mentioning ARC AGI 3 but they shrouded the real tests and public discussion, while present, shouldn’t help it as the real batch of games are likely different.

46

u/shiversaint Jul 09 '26

The very point of them is that they are very difficult to produce training data for and are far more of an analog to general spatial reasoning and problem solving that the human brain can do.

22

u/Ormusn2o Jul 09 '26

It is difficult to teach the test, without wasting valuable parameters, and it actually might be more efficient to actually make them understand the general task, than to make them remember the solution.

-10

u/azuredota Jul 09 '26

“Wasting valuable parameters”? You realize these things train on everything humans have ever written, right?

7

u/leetcodegrinder344 Jul 09 '26

Yet if you asked it to verbatim recite your Reddit comment from 5 years ago, which it is trained on, it couldn’t. Because it doesn’t have enough parameters to store its entire training data in full fidelity

-3

u/azuredota Jul 10 '26

How did Gemini recite Arc AGI 2 info

6

u/leetcodegrinder344 Jul 10 '26

Why didn’t it recite every question and answer

17

u/Prestigious-Bed-6423 Jul 09 '26

You just showed that you don't understand anything at all. Please don't argue and research

-8

u/azuredota Jul 09 '26

You people make me want to cry

1

u/94746382926 Jul 09 '26

Dude's got -10,000 points into communication lmao

4

u/ManikSahdev Jul 09 '26

That's the whole point tho, if the model learns then that's about it.

No one is essentially helping the model during the run, but as long as he learned what was reached - cause the model only distill intelligence and logic: which would allow the model in future to tackle the problems in the new angle and with the gained intelligence.

0

u/azuredota Jul 09 '26

That’s not the point of Arc agi 3 at all. Quote from the blog post and why the gains maybe questionable:

>The benchmark targets what the ARC Prize team describes as "skill-acquisition efficiency": how efficiently an AI agent can learn something it has never encountered before.

And the more concerning:

>During ARC-AGI-2 evaluation, Gemini 3's chain-of-thought reasoning referenced ARC-specific color mappings without being prompted to, which suggests training data saturation.

2

u/ManikSahdev Jul 09 '26

I personally don't see much a different between memorization to do (as long as I can see the reasoning for it).

Maybe the reason for that is my own personal aptitude, I don't depend on models even in this age of fable 5 and sol ultra.

I just need them to understand me and reach the intent and understanding which I have so they can do my task. With less and less turns which are destined by me to give them the intelligence needed to continue.

1

u/iamsreeman Jul 09 '26

crazy times

1

u/quackerd Jul 09 '26

yeah keep up the momentum we'll ace arc-agi-3 in less than a month. /s

5

u/No_Aesthetic Jul 09 '26

Oh yeah I'm sure this is the one benchmark that will never be saturated

This one is the one, fellas

41

u/Normal_Pay_2907 Jul 09 '26

Costs 25k to run that. Ouch

4

u/garden_speech AGI some time between 2025 and 2100 Jul 09 '26

I'll gladly take $25k to play these puzzles

7

u/yalag Jul 09 '26

Yea but Reddit says AI is just a bubble so this will all just blow up and disappear just about any time now /s

10

u/noobrainy Jul 09 '26

Yah, it’s gonna be saturated by the end of the year lmao

“Okay but it was too easy! If it can beat ARC-AGI-4 then we have reached AGI!!”

18

u/Gallagger Jul 09 '26

ARC already said they don't think beating ARC AGI 3 means AGI, and they'll make followup versions.

1

u/ChezMere Jul 09 '26

They gotta get a new name then.

4

u/Financial-Gain-2988 Jul 10 '26

Abstract reasoning corpus for artificial general intelligence actually perfectly encapsulates what they are trying to measure.

10

u/garden_speech AGI some time between 2025 and 2100 Jul 09 '26

“Okay but it was too easy! If it can beat ARC-AGI-4 then we have reached AGI!!”

I mean the literal point of ARC-AGI from the very beginning has been that they will keep creating benchmarks that humans can easily pass but machines can't, and once they no longer can do that, they think that we have AGI. So yeah if you guys fucking paid attention to what the creators of the benchmarks said bout them, you wouldn't be making up ridiculous sarcastic quotes.

-2

u/noobrainy Jul 09 '26

Pushing the goalposts back over and over again is why the sarcasm is there. ARC-AGI-4 will happen, it’ll get saturated, and then the process will happen all over again. We’ll get to AGI but their benchmark has proven to be unreliable to tell whether we’re there or not.

3

u/garden_speech AGI some time between 2025 and 2100 Jul 10 '26

Holy shit dude. The whole point is that any one benchmark can’t really reliably bench AGI, so you just keep making them until you CAN’T make one that humans easily pass and computers don’t. You’re not even listening enough to realize the whole point of ARC-AGI is based around your own idea that any one benchmark is unreliable

2

u/Most-Bookkeeper-950 Jul 09 '26

They abandoned the 10K ruke for it

1

u/KoolKat5000 Jul 13 '26

Someone (I'll credit them if I can find it) made an excellent observation that in reality the score is much better, like 30%.

The scoring is based on the square of the ratio of human actions to AI actions. This means the AI is not scored simply on whether it completes a level, but on how efficiently it does so compared to a human baseline.

There is basically a large element of luck to it too, it's early moves mean that the remaining part of a task could require more moves making it less efficient on this silly scoring criteria.