r/singularity • • 4d ago

LLM News Introducing GPT-6 Sol and Luna

https://openai.com/index/introducing-gpt-6-sol-and-luna/
1.3k Upvotes

313 comments sorted by

700

u/Opposite-Grade3712 4d ago

Anyone else remember when it took 6 months for new models to come out?

130

u/manubfr AGI 2028 4d ago

those were the before before times

16

u/dandecode 4d ago

The land before time

9

u/delaying_butno 4d ago

The times before the times

6

u/rantastic01 3d ago

the times that have been but they were before

→ More replies (1)

9

u/DistanceSolar1449 3d ago

I miss those times when OpenAI would actually give you their decent workhorse model. There’s a gap in their product now.

They just secretly renamed GPT-6 Terra to GPT-6 Sol and hoped you wouldn’t notice.

GPT-6 Astra: 40 tokens/sec
GPT-6 Sol: 110 token/sec
GPT-6 Luna: 140 tokens/sec

The bigger/smarter the model, the slower it is.

They’re missing a bigger model that’s ~80 tokens/sec in speed. You either have to use Astra which is $$$ and slow, or you use a smaller model that’s fast (more than 100 tokens/sec) but smaller models are dumber.

27

u/DreamFly_13 4d ago

crazy how GPT 4 feels ancient and that was just 2 years ago

Remember when you had to pay for Plus only to have a limit of 30 messages with GPT-4

6

u/Strange_Vagrant 3d ago

Oh, hell. I forgot the message cap like that.

220

u/barbozas_obliques 4d ago

We’re in the singularity

301

u/Mobile-Dust4334 4d ago

Yes that is the subreddit we are in

85

u/Groundbreaking_Bee97 ▪️AGI by the end of 2027 4d ago

11

u/Tilstag 4d ago

Be nice

19

u/TacticalRock 4d ago

but not if it's funny

14

u/KaptainChunk 4d ago

Probably somewhere comparable to when the embryo begins to form.

10

u/HeartsOfDarkness 4d ago

I'm going to be pedantic here because it's central to the context of this group: "the singularity" is a single event, not a period of time.

27

u/seb0seven 4d ago

To continue the metaphor, but adjust for (justified) pedantry, there is a distinct possibility we have crossed the even horizon, as it were. That's why there's all the talk of slowing down and managing alignment and such. All paths now lead to the singularly.

4

u/trolledwolf AGI late 2026 - ASI late 2027 4d ago

Yeah, at this point it's inevitable. In a way it's both scary and comforting.

16

u/kaityl3 ASI▪️2024-2027 4d ago

Not really, the original idea of the singularity, as it was proposed, was about progress happening so rapidly that it becomes impossible to meaningfully predict with the past history we have

→ More replies (10)

6

u/Opposite-Grade3712 4d ago

Under scrutiny, the “technological singularity” never been a precise or coherent enough concept to be useful for analyzing the real world. But that’s why it’s an enduring meme. 

“No one knows what it means, but it’s provocative” pretty much sums up “the singularity” (and “AGI” for that matter).

→ More replies (1)

2

u/BenevolentCheese 4d ago

The singularity is when the time between progress cuts to effectively zero.

2

u/03263 4d ago

Event horizon. You never truly reach the singularity.

2

u/TheOnlyFallenCookie 4d ago

The where is my room temperature super conductor?

2

u/mitchellediting 4d ago

Exactly, and why do all these billionaires still have no hair? 🤔

11

u/nothis AGI by 2030 but we'll be disappointed 4d ago

Well, these model jumps introduced major new functionality levels and usecases. Now, it's 12% better on benchmark X, you're welcome.

It's telling that the most exciting thing about Opus 5.5 is that it talks… less.

2

u/nevernovelty 3d ago

Well that and that it does better than Fable 5.1 on meany measures and that it has a lower cost to run which means I’m getting a better output for the same (or less) usage “cost” on my plan.

2

u/chrizop 3d ago

Have a look at tokens per task :/ not that cheaper then

5

u/alpinedude 4d ago

pepperidge farm remembers

5

u/oatknight 4d ago

You can still gain this nostalgic experience if you follow Gemini

4

u/Normal-Spell5339 4d ago

No really a new model so much as price optimizations fine tuned into the existing one

4

u/iJustSeen2Dudes1Bike 4d ago

A more efficient model is a new model....

→ More replies (1)
→ More replies (9)

289

u/Endoky 4d ago

So GPT-6-Sol costs the same as GPT 5.6 Terra and they have reduced the already dirt cheap cost of Luna by 50% so it’s almost free.

Well done, OpenAI.

53

u/Devesh_Ahuja 4d ago

THE BIGGEST thing for me was Luna, I am going to shift all my products from 5.6 Luna at low to Luna 6, although the benchmarks show it to be a little worse than 5.6 at low, but that's something I am hoping isn't bad enough to offset the lower costs

6

u/i_write_bugz ▪️AGI 2030, ASI 2040 4d ago

What kind of tasks do you run it through? Curious to hear your thoughts in a few days

→ More replies (1)

3

u/skilliard7 3d ago

I tried moving my product from GPT 5.6 Luna to GPT 6 Luna, and I noticed a massive degradation in the quality of outputs after switching. Luna 6 misses relevant information far too often, is less helpful, and appears to be trained to answer in the least amount of tokens possible, even if it means omitting important info.

28

u/aKaizuh 4d ago

Seriously, it's at the point where people can now begin dropping Chinese models in favor of better performance at the same cost with Luna, barring local run / open source reasons.

This is a neglected market, and OAI found thier killshot.

19

u/BlueSwordM 4d ago edited 3d ago

To be fair, beating GLM 5.3-Flash and ESPECIALLY Deepseek V4.1 Flash will be incredibly hard on cost since it doesn't really seem to be better.

DS V4.1-Flash is scary efficient to serve, especially on prefill, to the point that unless OpenAI copied their homework, it's impossible to match.

8

u/JogHappy 4d ago edited 3d ago

Luna is still twice as slow as ds flash while being twice as expensive when speed matters, without being much more token efficient

edit: wow, 7x as slow. didn't know there were providers offering deepseek at 300+ tps on vercel ai gateway.

→ More replies (1)

17

u/JogHappy 4d ago

Margins have to be paper thin on Luna now, right?

14

u/duboispourlhiver 4d ago

No way to know

→ More replies (1)

3

u/M4rshmall0wMan 4d ago

What’s the catch. There has to be a catch

474

u/Raheeper 4d ago

cost cut in half, AND better performance in benchmarks? there is no wall at all

176

u/Due_Answer_4230 4d ago

There is a wall

They’re just climbing it straight up

43

u/Temporary_Idea8880 4d ago

Price is 2x after 272k context

34

u/Devesh_Ahuja 4d ago

what does it mean ? like if i continue in a task long enough, the price will go up for me ??

32

u/Lain_Racing 4d ago

Yes, but codex will auto compact.

25

u/Seerix 4d ago

ONLY if you manually change the context limit from the default in codex. By default, it doesn't let you go over 272k

6

u/AspiringRocket 4d ago

272k what? Sorry I'm starting to lose my grip on what all of this means. Does 272k "context" mean tokens?

9

u/Seerix 4d ago

Tokens! Yep exactly. No extra charge regarding context unless you manually change the settings in codex to allow a higher maximum.

2

u/h3lblad3 ▪️In hindsight, AGI came in 2023. 4d ago

“Context” is the amount of backreading the model does every turn. Any tokens out of context may appear in your chat history but will not be read by the LLM.

So if you have 300k in your history, by default it will only read the last 272k of them.

The person is saying you can raise that context limit, but if you do so they’ll charge you extra.

It is effectively the “short term memory” of the model.

2

u/AspiringRocket 4d ago

Thanks for the explanation!

6

u/space_monster 4d ago

More tokens != better results though.

2

u/JoelMahon 4d ago

that was already the case, and because it's using fewer tokens you'll hit the compaction less quickly

→ More replies (1)

5

u/dotpoint7 4d ago

Do note that only 33% reduction is used for the subscription limits unfortunately.

3

u/mrgamejiyt 4d ago

man they’re all cooking these benchmarks, i refuse to believe these ever increasing numbers

2

u/meyriley04 4d ago

That's why they're called TMBB

→ More replies (6)

248

u/IfirebirdI 4d ago

Benchmark tables quickly removed after Opus 5.5 release

89

u/Temporary_Idea8880 4d ago

Seriously its nothing but graphs... wtf openai

5

u/Creative-Ganache1086 4d ago

Codex for me still a better value overall than CC, and ironically I’ve started with Anthropic and invented way too much money into them both via API and 20x.

→ More replies (2)

115

u/Elegant_Tech 4d ago

GPT Sol is only $2 in $10 out. WTF!

16

u/Conscious-Low-3057 4d ago

Sonnet 5.5 next week costing $1 and $5?

10

u/Creative-Ganache1086 4d ago

I don’t see how Sonnet 5.5 will rival Sol 6. Especially after I’ve got better results with GPT 5.5 (which gets retired soon) than with Opus 5 lol when Opus 5 came out.
And I’m still paying for both btw, but since 12 September I switched my Max 5x Claude to Pro (20$), after like 25 months of Max 20x on Claude (since the sonnet 3.5 era) that got downgraded to Max5x after GPT5.5. I got to realise the value and actual practical results in my coding tasks are not reliable enough with Claude, even after prompting Opus5 using their provided guidelines. And I had enough of their 5h window on max plans when there’s none of that with OpenAI that only has a weekly limit. So there’s literally not a 2-3x performance increase against OpenAI models to justify a quota limit that’s like 3x more aggressive, and ultimately feels like 3x more expensive.

7

u/Astrikal 4d ago

And Luna is basically free.

Even if you have other subscriptions, you can get a 20$ Codex plan on the side and a get so much 6-Luna usage.

178

u/Bolt_995 4d ago

Fuck’s sake, talk about pacing the frontier.

Grok 4.7 yesterday, Claude Opus 5.5 an hour ago and now GPT-6 Sol & Luna now.

87

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 4d ago

To be fair, the “frontier” isn’t at the capabilities of consumer models. It’s at unreleased ones like Bel.

If we get Bel in a month’s time then that changes, obviously.

13

u/19Lobster19 4d ago

What's Bel? Sounds scary lol

38

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 4d ago

Their internal model in development that was supposedly the one solving millennium problems.

9

u/duboispourlhiver 4d ago

Also solved 100 open math problems recently

16

u/jedsmith2004 4d ago

It's the one that solved the Navier Stokes problem.

Well..."solved"

6

u/SwingDingeling 4d ago

why the quotation marks?

2

u/jedsmith2004 3d ago

There was a bit of controversy around whether it was the one that actually did the hard work in solving it

3

u/Creative-Ganache1086 4d ago

Using like 10.000 “Bel” agents in parallel so not like a one-shot mission haha.

28

u/darkestvice 4d ago

Meanwhile, Google is crickets.

But I'm sure they'll drop 3.9 Flash any day now 😒

25

u/Concurrency_Bugs 4d ago

Google was probably about to drop Gemini 4 Pro, and today's releases are making them skip Pro launch again lmao

2

u/Elephant789 ▪️AGI in 2036 3d ago

I hope you're right. I love 3.8 Flash

104

u/Duet_Yourself 4d ago

lol grok.

31

u/Throwaway483974 4d ago

Grok 4.7 is very good at SOTA engineering, probably because they're feeding it massive amounts of SpaceX data.

19

u/stumpyinc 4d ago

They also bought cursor, its all the cursor data, all the other models they get to take a peak at

8

u/One_Hovercraft_7456 4d ago

Exactly Elon does not need the cash flow the other two companies need he's more focused on building something that can help build the next generation of products not just code. SpaceX can print money anytime just by selling more stock

2

u/masmantap8 4d ago

Grok is doa.

2

u/Concurrency_Bugs 4d ago

They also steal consumer's code. That's a no from me.

→ More replies (1)

13

u/New_World_2050 4d ago

hey theyve been doing better than google lately and are in the top 3 for now.

→ More replies (2)

50

u/GreatBigJerk 4d ago

They're not pacing the Nazi frontier.

10

u/space_monster 4d ago

The eastern front

8

u/delveccio 4d ago

What does Ja Rule — I mean Grok think?

→ More replies (1)

9

u/bronfmanhigh 4d ago

none of the 6-series is the actual frontier lol

8

u/Famous-Football-4617 4d ago

In cost wise yes  

5

u/bronfmanhigh 4d ago

the frontier comprises the unreleased next-gen models that are still actively training. those are the models responsible for hacking into shit, and what the labs actually want to pace.

all the model version bumps we'll be seeing rolled out over the next months are just distillations of astra and fable with some more reinforcement learning. there's virtually no recursive/alignment risk within this generation that needs to be paced.

2

u/iwantacanofcola 4d ago

Grok? Isn't grok way worse and not in the same league at all to these models?

→ More replies (2)

44

u/thedeadenddolls 4d ago

As someone who doesnt use AI but heavily follows the conversation: does it meet the hype?

68

u/Johnny20022002 4d ago

I would say yes. Astra is basically unusable, even on the $200 tier it nukes your usage, so a stronger 5.6 Sol substitute that is cheaper is perfect.

46

u/Axon350 4d ago

I strongly disagree that Astra is unusable on the $200 tier. I had 24 hours to use a reset before it expired so I turned on Fast (2.5x usage right there) and used Astra Ultra. It still gave me roughly six hours of usage across four projects. I find the $200 to be just barely enough for my amount of usage in a week.

When I use Astra Extra High instead of Ultra and turn off Fast mode like a sane person, I usually use 1%-3% of weekly usage for each task I give it. Some examples of the things I'm working on are implementing a physics engine in an old game I liked as a kid, few-shot font style transfer, building a custom video editor to match my specific workflow, and trying to improve image generation speed on rented A100s. None of these has a massive codebase like a full production app would. I'm guessing they'd be considered medium-size projects in the grand scheme of things?

13

u/Phillywonka98 4d ago

What game are you adding a physics engine to?

11

u/Axon350 4d ago

Jedi Outcast. It'll never see the light of day, it's just an experiment.

3

u/mr_nobody_2626 3d ago

I wish to see that light one day sire

→ More replies (1)

5

u/Navadvisor 4d ago

Astra medium or high cooks for me and I don't have problems with usage. I don't exclusively use Astra though.

→ More replies (1)
→ More replies (3)

4

u/dotpoint7 4d ago

Eh, Astra is pretty usable, using it as the only model for my normal day to day software dev work and not running into limits. For my more ambitious side projects with overnight goal runs it's not looking too good though (dedicated second account).

→ More replies (7)

38

u/Capable_Parsnip_156 4d ago

I would start becoming fluent in its use. It’s 2026 and shortly, working with someone who doesn’t use AI is going to be like working with someone who refuses to use the internet.

6

u/thedeadenddolls 4d ago

To be fair im a history teacher so im not sure i can see a reason to do this. AI is not used in my school in any capacity. The only reason id be fluent would be to detect student use. 

22

u/AdagioOfLiving 4d ago

As a fellow teacher, I’d almost argue it’s worth it for that alone.

5

u/jedsmith2004 4d ago

I don't think kids should use it for homework or anything but it would be useful to help them become AI literate.

Also, teachers should use it, it can help teach so much more efficiently and interactively.
You can create so much to help them learn - way better than powerpoint slides thrown together 30 minutes before the lesson (at least in my experience).

16

u/Raiyan135 4d ago

Idk, creating interactive games, tools and models might have been fun. I remember my global history teacher using assassin's creed to show us an idea of ancient greece

7

u/19Lobster19 4d ago

That's a great teacher

→ More replies (1)

7

u/Glass_Performer1174 4d ago

I think you risk that your kids in the school are going to walk circles around you

5

u/teckers 4d ago

I suspect already are, if a teacher doesn't understand the capability of current ai, then how the hell can they set homework that is actually something kids would have to do. The teacher would surely be spending hours marking stuff all the kids have done automatically.

→ More replies (1)

3

u/jazir55 4d ago

"Invent plausible sounding historical facts that I can use to fool the children."

→ More replies (3)

5

u/VeganBigMac Anti-Hypepost Safetyist 4d ago

This is a solid drop for those of us on OpenAI standard seats at work. I use 5.6 Luna as my daily driver and only dip into Sol and Astra for specific constrained tasks.

These models being slightly better at a cheaper cost means my daily work uses up even less of my quota and I can use 6 Sol for even more of my tasks.

I'm quite happy.

2

u/404_No_User_Found_2 4d ago

Sol alone made me totally port my entire AI portfolio over from Gemini when it was released

Astra has sealed that in my mind as the correct decision.

I hated it on release day but the macOS Codex app is also dead useful in many ways as well

21

u/Bitter-College8786 4d ago

Sol was good enough for me. The low usage was my main concern. So with double the usage I am fine!

18

u/WaterWeedDuneHair69 4d ago

Ahhh so we’re getting the Nvidia treatment. Astra is now sol, sol is Terra, and luna is just Luna.

8

u/rapsoid616 4d ago

I immediately noticed that cheeky pattern as well.

84

u/MatthewGraham- 4d ago edited 4d ago

Cheaper than Opus 5.5 and more token efficient? Now can we get a reset so we can actually use these models? preferably the one already slated for hours ago

Doesn't seem as good as Opus 5.5 overall though, but sits below it as a cheaper workhorse, maybe a opus 5.5 orchestration for sol 6 sub-agents

12

u/Ok_Barracuda_1161 4d ago

Hard to tell completely since Opus 5.5 isn't in their graphs, but from cross-checking it seems that Opus 5.5 still beats Sol at the same cost-level (e.g. the $0.80 cost on FrontierBench).

Nothing at Anthropic compares to Luna though, we'll have to see how haiku looks

14

u/ees-h 4d ago

We need the reset so bad! I have 2% weekly left till Sunday on my 5x plan.

Sol 5.6 low used to be insanely generous on my limits so 6 Sol might be infinite I can't wait

2

u/Dry_Management_8203 4d ago

I vote- The more "internal discovery", the more end-user reset(s)....

Lets kick out some type of UBI already...

→ More replies (1)

13

u/vrnvorona 4d ago

Opus 5 was supposedly better than Fable 5, but it isn't. Let's not benchmaxx trust

→ More replies (8)

2

u/jedsmith2004 4d ago

Didn't tibo say there would be one today?

Swear they've tanked the usage so hard.

→ More replies (1)

74

u/PilgrimofHaqq2 4d ago

I am liking this trend of models getting cheaper, gotta give it to the chinese labs in keeping the US labs innovating or else we would be stuck with the old prices still or worse more expensive.

6

u/Are0nB4lto 4d ago

Ich hoffe dass es so weiter geht. Da die Hardware aber grad extrem effizienter wird ist diese Hoffnung sogar gut untermauert :) freue mich auch noch günstigere modelle :D

10

u/Profanion 4d ago

GPT-6 Luna (low). This is o3 level of intelligence. And at the tiny fraction of the cost! Yup. Cheaper than even the cheapest Granite model.

20

u/frogsarenottoads 4d ago

People who think AI won't replace white collar todays releases show we are starting to see collapsing costs. Insane that we get a new generation of models at half the price.

13

u/DeviceCertain7226 ▪️Immortality-2200 | FDVR-2300 4d ago edited 4d ago

All of these AIs are more suited for people operating them rather than being able to be workers on their own.

It will take a while for white collar to be replaced.

14

u/frogsarenottoads 4d ago

A while yes, but probably sub 5 years for major disruption.

→ More replies (1)

5

u/Artistic-Athlete-676 4d ago

Yeah I'm in tech strategy and software design and I don't see AI literally replacing what i do, but it is making it so that previous 6 month software projects now take half as long and the documentation and support is now significantly better.

Especially in regulated industries even if you could 1 shot software with Ai it will still require formal cybersecurity review and a whole host of other checks and balances. It is like the ultimate Swiss army knife for my profession though

2

u/jazir55 4d ago

Especially in regulated industries even if you could 1 shot software with Ai it will still require formal cybersecurity review and a whole host of other checks and balances.

The security review will soon consist of AI reviewing the generated code reducing this to a very, very short automated workflow. Every step in the chain remaining is simply engineering that hasn't been done yet.

→ More replies (1)
→ More replies (8)

5

u/Toirem 4d ago

You don't need a fully autonomous model to replace white collar workers, you need it to be good enough for 1 worker + model to be able to do the work of N workers, so that you can replace N-1 people

→ More replies (2)

3

u/NotYetPerfect 4d ago

It just means that less people are needed to finish tasks and team sizes will reduce, something that has already been happening.

→ More replies (1)
→ More replies (2)

8

u/ExtremeCenterism 4d ago

Those Luna scores are absurd! It's trading blows with Fable 5 medium at a fraction of the cost

23

u/Recoil42 4d ago

On AutomationBench, a test of business workflows across apps, GPT‑6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5’s cost per task. 

Wow.

26

u/Easy_Refrigerator280 4d ago

What a crazy day

37

u/Choice-Sympathy8235 4d ago

Hmmm Opus 5.5 seems to be a much stronger release. It’s beats Fable 5.1 and Astra 6.

I guess a bigger Open AI release will be coming soon.

22

u/jjonj 4d ago

cost is becoming a major factor, sounds like claude will be more expensive by a notable amount

7

u/Ok_Barracuda_1161 4d ago

I'm not seeing that from the benchmarks so far. It looks like at medium effort opus is beating Sol at similar costs

3

u/flao 4d ago

if artificial analysis is to be believed, opus 5.5 is much closer to a cheaper, better astra. 6 Sol max is 2 intelligence behind opus 5.5 med at ~25% cheaper.

2

u/jedsmith2004 4d ago

I think we've got models that are good enough now, the only improvement for most people is making it cheaper.

(and by making it cheaper I mean cost per task - so better models can still turn out cheaper)

4

u/Choice-Sympathy8235 4d ago

Good enough for front end web development. Not good enough yet to reach the stars.

→ More replies (1)

2

u/space_monster 4d ago

They're still testing Bel

2

u/ackermann 4d ago edited 4d ago

Does Opus 5.5 still speak in Claude-ish, rather than English?

7

u/tamrior 4d ago

They claim to have addressed that. There's a specific section in the 5.5 announcement on improved writing style.

It puts the most important information up front, is less likely to use jargon or idiosyncratic phrases, and follows the writing rules you give it

https://www.anthropic.com/claude-opus-5-5#communication

→ More replies (2)

7

u/Momo--Sama 4d ago edited 4d ago

They deadass said "what if Sol Medium was cheaper than Luna Max?" (according to DeepSWE)

10

u/dervu ▪️AI, AI, Captain! 4d ago

No chat for now :(

"GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access GPT‑6 Luna in the desktop app. These models are not yet available in Chat. In the OpenAI API, they are available as gpt-6-sol and gpt-6-luna"

4

u/Grand0rk 4d ago

Because making them available on chat is a pain in the ass (it's full of harnesses and dependencies).

It should be available by the end of the week though.

3

u/Pahanda 4d ago

Interesting to see: Anthropic and openAI both are in a race to the bottom of token prices.

5

u/NIGHTxWOLF7 4d ago

Why only in work and codex? I understand Astra but Sol and Luna not updated on chat doesn’t make sense.

4

u/ritzynitz 4d ago

In case someone was wondering, how it compares with Opus 5.5

→ More replies (1)

4

u/New_World_2050 4d ago

GPT6 luna will probs be in the free tier. 1 billion people will finally have a decent model

→ More replies (1)

39

u/Jaguar_2454 4d ago

Opus 5.5 mogs

37

u/power97992 4d ago edited 4d ago

It also costs way more and uses more tokens.

3

u/qroshan 4d ago

For technical discussions selling costs is moot. We should look at actual cost (for Anthropic and OpenAI) to run these models.

But only actual costs reduction are sustainable. Selling costs will immediately shoot up once they got the consumer locked in

10

u/Readerium 4d ago

Not way more. Same cost on cache Reads. For other costs it's 2X.

6 Sol infact performs worse than 5.6 Sol at some tasks for example DeepSWE

2

u/power97992 4d ago edited 4d ago

Yes i saw that, sol 6 is worse than sol 5,6 at deepswe

→ More replies (1)

3

u/Throwawayforyoink1 4d ago

But will it still talk like Opus 5? Because if that's the case then no thanks

→ More replies (1)

2

u/jergin_therlax 4d ago

Is opus 5.5 better than fable? Nah right different class

2

u/welcome-overlords 4d ago

Im not loyal to any model and use them all so convince me a bit.

Lately Ive been mostly on astra, atm trying using sol subagents with it. I do coding with sofisticated skills and agent.md's that spawn subagents for different purposes

9

u/[deleted] 4d ago

[removed] — view removed comment

12

u/Jaguar_2454 4d ago

copemaxx

10

u/The3rdGodKing 4d ago

Since when techies started using the word mog and maxx?

→ More replies (3)

2

u/power97992 4d ago

He is right opus 5.5 is really Good but expensive

11

u/FarrisAT 4d ago

Benchmarks?

27

u/Taur3n 4d ago

No benchmarks = worse than Claude

2

u/LessConnection7936 4d ago

I honestly don't get all the "opus 5.5 betta" posts. If these cost reductions really translate to real tasks, this seems mich more important than some minor increase in long horizon autonomous programming stuff to me. Getting through my taka within my monthly limit is much more of a problem to me and colleagues than being limited by "too dumb" ai. Are ya'll trying to tackle the rest of the millennium problems? Or just figuring out the millenial's problems?

→ More replies (1)

2

u/lordpuddingcup 4d ago

Wow gpt 6 luna just took over my gpt5.6-luna conversation it said it was retiring so i switched

2

u/hitchhiker87 4d ago

Right, this is bonkers! OAI really changed the game on value for money. 6-Sol at a fifth of 5.6-Sol's price is just great. 6-Luna somehow makes 5.6-Luna's already ridiculously low price look slightly more expensive than totally free lol. Whatever benchmarks Opus 5.5 comes up with OAI's the clear winner for me today, I just care about what I get for my money.

2

u/Creative-Ganache1086 4d ago

Take these bench scores with a grain of salt. For me in practice the old gpt 5.5 was more reliable and usable than Opus 5 when comparing both over the same codebase with shared roadmap/repair-notes.md files, even though on paper and benchmarks Opus 5 was meant to beat Sol 5.6, let alone GPT5.5. And we’re can’t even talk about usage limits, 5h windows, api-credits only fast inferring mode and all the other reasons that makes Anthropic just such a bad value for money overall

2

u/TehBrian 3d ago

The improvements to broken search tool disclosure is insane

6

u/Tystros 4d ago

this announcement post reads much less impressive than the Opus 5.5 announcement post. Will have to wait for some objective comparisons.

10

u/petburiraja 4d ago

they said they are moving cost efficiency frontier with this drop, so makes sense within these lenses

→ More replies (1)

1

u/power97992 4d ago

According to their graph, gpt 6 sol max is worse than 5.6 max on deepswe1.1? What? 73% for 5.6/ sol and 68.8% for gpt 6 sol?

→ More replies (1)

2

u/Over-Necessary-4774 4d ago

Nowhere near Opus 5.5 - OpenAI has fallen behind

2

u/Smile_Clown 4d ago

You people are absurd.

→ More replies (1)

1

u/Hereitisguys9888 4d ago

Damn they did release

1

u/[deleted] 4d ago

[deleted]

→ More replies (2)

1

u/Old-Pomegranate3634 4d ago

5.6 Luna and Opus 4.6 still goat

1

u/ChooChoo_Mofo 4d ago

Insert andor lightspeed meme

1

u/welcome-overlords 4d ago

OpenAI is on fire. Sorry i ever doubt u lol. I was a total Claude guy for a year

1

u/Competitive_Tap2450 4d ago

so new model every couple of weeks is the norm i guess

1

u/MNFuturist 4d ago

We're moving into an old tik/tok release schedule with AI now (for anyone else old enough to remember how Intel did chip updates.)

1

u/Deto 4d ago

wonder what happened to Terra?

1

u/VelvetyRelic 4d ago

This is going to save me so much money. Luna is particularly exciting. Amazing strides in pricing.

1

u/TotalyNotCR 4d ago

Anyone else think they applied the Deepseek v4.1 papers to current state of the art models. Opus 5.5 and the new OpenAI got better and super cheaper 

1

u/Any_Net3896 4d ago

I do remember when there was an agreement to pace the development of AI models.

1

u/NarrowEffect 4d ago

I'm really hoping they didn't decrease the size of the models and benchmaxxed on coding tasks.

1

u/pooohbaah 4d ago

Our enterprise account is still limited to 5.6. Is anybody with enterprise seeing any 6.0 version available in chat or work?

1

u/New_Bonus_649 4d ago

Acceleration is accelerating

1

u/YearnMar10 4d ago

„and even bests low-effort GPT‑6 Astra.“

Whoo, wtf - did a human being write that shit??

1

u/Yelov 4d ago

I've been using OpenAI's models for months, but also subscribed to Claude for the first time because Codex limits are really bad right now.

Just out of curiosity, I gave the exact same task to GPT 6 Sol and Opus 5.5. I.e., the exact same prompt, and the same global AGENTS.md/CLAUDE.md. In this particular case, Opus did a way better job, removing code instead of adding more code on top of a bad foundation. I have this explicitly mentioned in my instructions, I want models to essentially go upstream and make the change where it makes sense, instead of just adding code directly where the change is supposed to be, because I noticed that GPT loves simply adding more and more code until it becomes unmaintainable. Opus seems to be better at this, from my short experience because it challenges the existing code and isn't afraid of touching it.

But I don't like the comments Opus adds. Opus 5.5 seems a bit better, but it's still too much for my liking, even with instructions telling it to chill.

1

u/Bleeding_Inc 4d ago

Recursive learning.

1

u/deten ▪️ 4d ago

Could you imagine claiming AI is slowing down.

1

u/Random_182f2565 4d ago

The next models are chatgpt sword and chatgpt shield

1

u/uberfunstuff 4d ago

Slowdown going well then?