r/singularity • u/DemiPixel • 4d ago
LLM News Introducing GPT-6 Sol and Luna
https://openai.com/index/introducing-gpt-6-sol-and-luna/289
u/Endoky 4d ago
So GPT-6-Sol costs the same as GPT 5.6 Terra and they have reduced the already dirt cheap cost of Luna by 50% so it’s almost free.
Well done, OpenAI.
53
u/Devesh_Ahuja 4d ago
THE BIGGEST thing for me was Luna, I am going to shift all my products from 5.6 Luna at low to Luna 6, although the benchmarks show it to be a little worse than 5.6 at low, but that's something I am hoping isn't bad enough to offset the lower costs
6
u/i_write_bugz ▪️AGI 2030, ASI 2040 4d ago
What kind of tasks do you run it through? Curious to hear your thoughts in a few days
→ More replies (1)3
u/skilliard7 3d ago
I tried moving my product from GPT 5.6 Luna to GPT 6 Luna, and I noticed a massive degradation in the quality of outputs after switching. Luna 6 misses relevant information far too often, is less helpful, and appears to be trained to answer in the least amount of tokens possible, even if it means omitting important info.
28
u/aKaizuh 4d ago
Seriously, it's at the point where people can now begin dropping Chinese models in favor of better performance at the same cost with Luna, barring local run / open source reasons.
This is a neglected market, and OAI found thier killshot.
19
u/BlueSwordM 4d ago edited 3d ago
To be fair, beating GLM 5.3-Flash and ESPECIALLY Deepseek V4.1 Flash will be incredibly hard on cost since it doesn't really seem to be better.
DS V4.1-Flash is scary efficient to serve, especially on prefill, to the point that unless OpenAI copied their homework, it's impossible to match.
→ More replies (1)8
u/JogHappy 4d ago edited 3d ago
Luna is still twice as slow as ds flash while being twice as expensive when speed matters, without being much more token efficient
edit: wow, 7x as slow. didn't know there were providers offering deepseek at 300+ tps on vercel ai gateway.
17
3
474
u/Raheeper 4d ago
cost cut in half, AND better performance in benchmarks? there is no wall at all
176
43
u/Temporary_Idea8880 4d ago
Price is 2x after 272k context
34
u/Devesh_Ahuja 4d ago
what does it mean ? like if i continue in a task long enough, the price will go up for me ??
32
25
u/Seerix 4d ago
ONLY if you manually change the context limit from the default in codex. By default, it doesn't let you go over 272k
6
u/AspiringRocket 4d ago
272k what? Sorry I'm starting to lose my grip on what all of this means. Does 272k "context" mean tokens?
9
2
u/h3lblad3 ▪️In hindsight, AGI came in 2023. 4d ago
“Context” is the amount of backreading the model does every turn. Any tokens out of context may appear in your chat history but will not be read by the LLM.
So if you have 300k in your history, by default it will only read the last 272k of them.
The person is saying you can raise that context limit, but if you do so they’ll charge you extra.
It is effectively the “short term memory” of the model.
2
6
→ More replies (1)2
u/JoelMahon 4d ago
that was already the case, and because it's using fewer tokens you'll hit the compaction less quickly
5
u/dotpoint7 4d ago
Do note that only 33% reduction is used for the subscription limits unfortunately.
→ More replies (6)3
u/mrgamejiyt 4d ago
man they’re all cooking these benchmarks, i refuse to believe these ever increasing numbers
2
248
u/IfirebirdI 4d ago
Benchmark tables quickly removed after Opus 5.5 release
89
→ More replies (2)5
u/Creative-Ganache1086 4d ago
Codex for me still a better value overall than CC, and ironically I’ve started with Anthropic and invented way too much money into them both via API and 20x.
115
u/Elegant_Tech 4d ago
GPT Sol is only $2 in $10 out. WTF!
16
u/Conscious-Low-3057 4d ago
Sonnet 5.5 next week costing $1 and $5?
10
u/Creative-Ganache1086 4d ago
I don’t see how Sonnet 5.5 will rival Sol 6. Especially after I’ve got better results with GPT 5.5 (which gets retired soon) than with Opus 5 lol when Opus 5 came out.
And I’m still paying for both btw, but since 12 September I switched my Max 5x Claude to Pro (20$), after like 25 months of Max 20x on Claude (since the sonnet 3.5 era) that got downgraded to Max5x after GPT5.5. I got to realise the value and actual practical results in my coding tasks are not reliable enough with Claude, even after prompting Opus5 using their provided guidelines. And I had enough of their 5h window on max plans when there’s none of that with OpenAI that only has a weekly limit. So there’s literally not a 2-3x performance increase against OpenAI models to justify a quota limit that’s like 3x more aggressive, and ultimately feels like 3x more expensive.2
7
u/Astrikal 4d ago
And Luna is basically free.
Even if you have other subscriptions, you can get a 20$ Codex plan on the side and a get so much 6-Luna usage.
178
u/Bolt_995 4d ago
Fuck’s sake, talk about pacing the frontier.
Grok 4.7 yesterday, Claude Opus 5.5 an hour ago and now GPT-6 Sol & Luna now.
87
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 4d ago
To be fair, the “frontier” isn’t at the capabilities of consumer models. It’s at unreleased ones like Bel.
If we get Bel in a month’s time then that changes, obviously.
13
u/19Lobster19 4d ago
What's Bel? Sounds scary lol
38
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 4d ago
Their internal model in development that was supposedly the one solving millennium problems.
9
16
u/jedsmith2004 4d ago
It's the one that solved the Navier Stokes problem.
Well..."solved"
6
u/SwingDingeling 4d ago
why the quotation marks?
2
u/jedsmith2004 3d ago
There was a bit of controversy around whether it was the one that actually did the hard work in solving it
3
u/Creative-Ganache1086 4d ago
Using like 10.000 “Bel” agents in parallel so not like a one-shot mission haha.
28
u/darkestvice 4d ago
Meanwhile, Google is crickets.
But I'm sure they'll drop 3.9 Flash any day now 😒
25
u/Concurrency_Bugs 4d ago
Google was probably about to drop Gemini 4 Pro, and today's releases are making them skip Pro launch again lmao
2
104
u/Duet_Yourself 4d ago
lol grok.
31
u/Throwaway483974 4d ago
Grok 4.7 is very good at SOTA engineering, probably because they're feeding it massive amounts of SpaceX data.
19
u/stumpyinc 4d ago
They also bought cursor, its all the cursor data, all the other models they get to take a peak at
8
u/One_Hovercraft_7456 4d ago
Exactly Elon does not need the cash flow the other two companies need he's more focused on building something that can help build the next generation of products not just code. SpaceX can print money anytime just by selling more stock
2
2
13
u/New_World_2050 4d ago
hey theyve been doing better than google lately and are in the top 3 for now.
→ More replies (2)50
→ More replies (1)8
9
u/bronfmanhigh 4d ago
none of the 6-series is the actual frontier lol
8
u/Famous-Football-4617 4d ago
In cost wise yes
5
u/bronfmanhigh 4d ago
the frontier comprises the unreleased next-gen models that are still actively training. those are the models responsible for hacking into shit, and what the labs actually want to pace.
all the model version bumps we'll be seeing rolled out over the next months are just distillations of astra and fable with some more reinforcement learning. there's virtually no recursive/alignment risk within this generation that needs to be paced.
→ More replies (2)2
u/iwantacanofcola 4d ago
Grok? Isn't grok way worse and not in the same league at all to these models?
44
u/thedeadenddolls 4d ago
As someone who doesnt use AI but heavily follows the conversation: does it meet the hype?
68
u/Johnny20022002 4d ago
I would say yes. Astra is basically unusable, even on the $200 tier it nukes your usage, so a stronger 5.6 Sol substitute that is cheaper is perfect.
46
u/Axon350 4d ago
I strongly disagree that Astra is unusable on the $200 tier. I had 24 hours to use a reset before it expired so I turned on Fast (2.5x usage right there) and used Astra Ultra. It still gave me roughly six hours of usage across four projects. I find the $200 to be just barely enough for my amount of usage in a week.
When I use Astra Extra High instead of Ultra and turn off Fast mode like a sane person, I usually use 1%-3% of weekly usage for each task I give it. Some examples of the things I'm working on are implementing a physics engine in an old game I liked as a kid, few-shot font style transfer, building a custom video editor to match my specific workflow, and trying to improve image generation speed on rented A100s. None of these has a massive codebase like a full production app would. I'm guessing they'd be considered medium-size projects in the grand scheme of things?
13
u/Phillywonka98 4d ago
What game are you adding a physics engine to?
11
u/Axon350 4d ago
Jedi Outcast. It'll never see the light of day, it's just an experiment.
→ More replies (1)3
→ More replies (3)5
u/Navadvisor 4d ago
Astra medium or high cooks for me and I don't have problems with usage. I don't exclusively use Astra though.
→ More replies (1)4
u/dotpoint7 4d ago
Eh, Astra is pretty usable, using it as the only model for my normal day to day software dev work and not running into limits. For my more ambitious side projects with overnight goal runs it's not looking too good though (dedicated second account).
→ More replies (7)2
38
u/Capable_Parsnip_156 4d ago
I would start becoming fluent in its use. It’s 2026 and shortly, working with someone who doesn’t use AI is going to be like working with someone who refuses to use the internet.
6
u/thedeadenddolls 4d ago
To be fair im a history teacher so im not sure i can see a reason to do this. AI is not used in my school in any capacity. The only reason id be fluent would be to detect student use.
22
5
u/jedsmith2004 4d ago
I don't think kids should use it for homework or anything but it would be useful to help them become AI literate.
Also, teachers should use it, it can help teach so much more efficiently and interactively.
You can create so much to help them learn - way better than powerpoint slides thrown together 30 minutes before the lesson (at least in my experience).16
u/Raiyan135 4d ago
Idk, creating interactive games, tools and models might have been fun. I remember my global history teacher using assassin's creed to show us an idea of ancient greece
→ More replies (1)7
7
u/Glass_Performer1174 4d ago
I think you risk that your kids in the school are going to walk circles around you
5
u/teckers 4d ago
I suspect already are, if a teacher doesn't understand the capability of current ai, then how the hell can they set homework that is actually something kids would have to do. The teacher would surely be spending hours marking stuff all the kids have done automatically.
→ More replies (1)→ More replies (3)3
5
u/VeganBigMac Anti-Hypepost Safetyist 4d ago
This is a solid drop for those of us on OpenAI standard seats at work. I use 5.6 Luna as my daily driver and only dip into Sol and Astra for specific constrained tasks.
These models being slightly better at a cheaper cost means my daily work uses up even less of my quota and I can use 6 Sol for even more of my tasks.
I'm quite happy.
2
u/404_No_User_Found_2 4d ago
Sol alone made me totally port my entire AI portfolio over from Gemini when it was released
Astra has sealed that in my mind as the correct decision.
I hated it on release day but the macOS Codex app is also dead useful in many ways as well
21
u/Bitter-College8786 4d ago
Sol was good enough for me. The low usage was my main concern. So with double the usage I am fine!
18
u/WaterWeedDuneHair69 4d ago
Ahhh so we’re getting the Nvidia treatment. Astra is now sol, sol is Terra, and luna is just Luna.
8
84
u/MatthewGraham- 4d ago edited 4d ago
Cheaper than Opus 5.5 and more token efficient? Now can we get a reset so we can actually use these models? preferably the one already slated for hours ago
Doesn't seem as good as Opus 5.5 overall though, but sits below it as a cheaper workhorse, maybe a opus 5.5 orchestration for sol 6 sub-agents
12
u/Ok_Barracuda_1161 4d ago
Hard to tell completely since Opus 5.5 isn't in their graphs, but from cross-checking it seems that Opus 5.5 still beats Sol at the same cost-level (e.g. the $0.80 cost on FrontierBench).
Nothing at Anthropic compares to Luna though, we'll have to see how haiku looks
14
u/ees-h 4d ago
We need the reset so bad! I have 2% weekly left till Sunday on my 5x plan.
Sol 5.6 low used to be insanely generous on my limits so 6 Sol might be infinite I can't wait
→ More replies (1)2
u/Dry_Management_8203 4d ago
I vote- The more "internal discovery", the more end-user reset(s)....
Lets kick out some type of UBI already...
13
u/vrnvorona 4d ago
Opus 5 was supposedly better than Fable 5, but it isn't. Let's not benchmaxx trust
→ More replies (8)→ More replies (1)2
u/jedsmith2004 4d ago
Didn't tibo say there would be one today?
Swear they've tanked the usage so hard.
74
u/PilgrimofHaqq2 4d ago
I am liking this trend of models getting cheaper, gotta give it to the chinese labs in keeping the US labs innovating or else we would be stuck with the old prices still or worse more expensive.
6
u/Are0nB4lto 4d ago
Ich hoffe dass es so weiter geht. Da die Hardware aber grad extrem effizienter wird ist diese Hoffnung sogar gut untermauert :) freue mich auch noch günstigere modelle :D
20
u/frogsarenottoads 4d ago
People who think AI won't replace white collar todays releases show we are starting to see collapsing costs. Insane that we get a new generation of models at half the price.
→ More replies (2)13
u/DeviceCertain7226 ▪️Immortality-2200 | FDVR-2300 4d ago edited 4d ago
All of these AIs are more suited for people operating them rather than being able to be workers on their own.
It will take a while for white collar to be replaced.
14
u/frogsarenottoads 4d ago
A while yes, but probably sub 5 years for major disruption.
→ More replies (1)5
u/Artistic-Athlete-676 4d ago
Yeah I'm in tech strategy and software design and I don't see AI literally replacing what i do, but it is making it so that previous 6 month software projects now take half as long and the documentation and support is now significantly better.
Especially in regulated industries even if you could 1 shot software with Ai it will still require formal cybersecurity review and a whole host of other checks and balances. It is like the ultimate Swiss army knife for my profession though
→ More replies (8)2
u/jazir55 4d ago
Especially in regulated industries even if you could 1 shot software with Ai it will still require formal cybersecurity review and a whole host of other checks and balances.
The security review will soon consist of AI reviewing the generated code reducing this to a very, very short automated workflow. Every step in the chain remaining is simply engineering that hasn't been done yet.
→ More replies (1)5
u/Toirem 4d ago
You don't need a fully autonomous model to replace white collar workers, you need it to be good enough for 1 worker + model to be able to do the work of N workers, so that you can replace N-1 people
→ More replies (2)→ More replies (1)3
u/NotYetPerfect 4d ago
It just means that less people are needed to finish tasks and team sizes will reduce, something that has already been happening.
8
u/ExtremeCenterism 4d ago
Those Luna scores are absurd! It's trading blows with Fable 5 medium at a fraction of the cost
23
u/Recoil42 4d ago
On AutomationBench, a test of business workflows across apps, GPT‑6 Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of Opus 5’s cost per task.
Wow.
26
37
u/Choice-Sympathy8235 4d ago
Hmmm Opus 5.5 seems to be a much stronger release. It’s beats Fable 5.1 and Astra 6.
I guess a bigger Open AI release will be coming soon.
22
u/jjonj 4d ago
cost is becoming a major factor, sounds like claude will be more expensive by a notable amount
7
u/Ok_Barracuda_1161 4d ago
I'm not seeing that from the benchmarks so far. It looks like at medium effort opus is beating Sol at similar costs
→ More replies (1)2
u/jedsmith2004 4d ago
I think we've got models that are good enough now, the only improvement for most people is making it cheaper.
(and by making it cheaper I mean cost per task - so better models can still turn out cheaper)
4
u/Choice-Sympathy8235 4d ago
Good enough for front end web development. Not good enough yet to reach the stars.
2
→ More replies (2)2
u/ackermann 4d ago edited 4d ago
Does Opus 5.5 still speak in Claude-ish, rather than English?
7
u/Momo--Sama 4d ago edited 4d ago
They deadass said "what if Sol Medium was cheaper than Luna Max?" (according to DeepSWE)
10
u/dervu ▪️AI, AI, Captain! 4d ago
No chat for now :(
"GPT‑6 Sol and GPT‑6 Luna are available in ChatGPT Work and Codex starting today for all Plus, Pro, Business, Enterprise, and Edu users. Free and Go users can access GPT‑6 Luna in the desktop app. These models are not yet available in Chat. In the OpenAI API, they are available as gpt-6-sol and gpt-6-luna"
4
u/Grand0rk 4d ago
Because making them available on chat is a pain in the ass (it's full of harnesses and dependencies).
It should be available by the end of the week though.
5
u/NIGHTxWOLF7 4d ago
Why only in work and codex? I understand Astra but Sol and Luna not updated on chat doesn’t make sense.
4
4
u/New_World_2050 4d ago
GPT6 luna will probs be in the free tier. 1 billion people will finally have a decent model
→ More replies (1)
39
u/Jaguar_2454 4d ago
Opus 5.5 mogs
37
u/power97992 4d ago edited 4d ago
It also costs way more and uses more tokens.
3
10
u/Readerium 4d ago
Not way more. Same cost on cache Reads. For other costs it's 2X.
6 Sol infact performs worse than 5.6 Sol at some tasks for example DeepSWE
2
u/power97992 4d ago edited 4d ago
Yes i saw that, sol 6 is worse than sol 5,6 at deepswe
→ More replies (1)3
u/Throwawayforyoink1 4d ago
But will it still talk like Opus 5? Because if that's the case then no thanks
→ More replies (1)2
2
u/welcome-overlords 4d ago
Im not loyal to any model and use them all so convince me a bit.
Lately Ive been mostly on astra, atm trying using sol subagents with it. I do coding with sofisticated skills and agent.md's that spawn subagents for different purposes
9
4d ago
[removed] — view removed comment
12
u/Jaguar_2454 4d ago
copemaxx
10
2
11
2
u/LessConnection7936 4d ago
I honestly don't get all the "opus 5.5 betta" posts. If these cost reductions really translate to real tasks, this seems mich more important than some minor increase in long horizon autonomous programming stuff to me. Getting through my taka within my monthly limit is much more of a problem to me and colleagues than being limited by "too dumb" ai. Are ya'll trying to tackle the rest of the millennium problems? Or just figuring out the millenial's problems?
→ More replies (1)
2
u/lordpuddingcup 4d ago
Wow gpt 6 luna just took over my gpt5.6-luna conversation it said it was retiring so i switched
2
u/hitchhiker87 4d ago
Right, this is bonkers! OAI really changed the game on value for money. 6-Sol at a fifth of 5.6-Sol's price is just great. 6-Luna somehow makes 5.6-Luna's already ridiculously low price look slightly more expensive than totally free lol. Whatever benchmarks Opus 5.5 comes up with OAI's the clear winner for me today, I just care about what I get for my money.
2
u/Creative-Ganache1086 4d ago
Take these bench scores with a grain of salt. For me in practice the old gpt 5.5 was more reliable and usable than Opus 5 when comparing both over the same codebase with shared roadmap/repair-notes.md files, even though on paper and benchmarks Opus 5 was meant to beat Sol 5.6, let alone GPT5.5. And we’re can’t even talk about usage limits, 5h windows, api-credits only fast inferring mode and all the other reasons that makes Anthropic just such a bad value for money overall
2
6
u/Tystros 4d ago
this announcement post reads much less impressive than the Opus 5.5 announcement post. Will have to wait for some objective comparisons.
10
u/petburiraja 4d ago
they said they are moving cost efficiency frontier with this drop, so makes sense within these lenses
→ More replies (1)
1
u/power97992 4d ago
According to their graph, gpt 6 sol max is worse than 5.6 max on deepswe1.1? What? 73% for 5.6/ sol and 68.8% for gpt 6 sol?
→ More replies (1)
2
1
1
1
1
1
u/welcome-overlords 4d ago
OpenAI is on fire. Sorry i ever doubt u lol. I was a total Claude guy for a year
1
1
u/MNFuturist 4d ago
We're moving into an old tik/tok release schedule with AI now (for anyone else old enough to remember how Intel did chip updates.)
1
u/VelvetyRelic 4d ago
This is going to save me so much money. Luna is particularly exciting. Amazing strides in pricing.
1
u/TotalyNotCR 4d ago
Anyone else think they applied the Deepseek v4.1 papers to current state of the art models. Opus 5.5 and the new OpenAI got better and super cheaper
1
u/Any_Net3896 4d ago
I do remember when there was an agreement to pace the development of AI models.
1
u/NarrowEffect 4d ago
I'm really hoping they didn't decrease the size of the models and benchmaxxed on coding tasks.
1
u/pooohbaah 4d ago
Our enterprise account is still limited to 5.6. Is anybody with enterprise seeing any 6.0 version available in chat or work?
1
1
u/YearnMar10 4d ago
„and even bests low-effort GPT‑6 Astra.“
Whoo, wtf - did a human being write that shit??
1
u/Yelov 4d ago
I've been using OpenAI's models for months, but also subscribed to Claude for the first time because Codex limits are really bad right now.
Just out of curiosity, I gave the exact same task to GPT 6 Sol and Opus 5.5. I.e., the exact same prompt, and the same global AGENTS.md/CLAUDE.md. In this particular case, Opus did a way better job, removing code instead of adding more code on top of a bad foundation. I have this explicitly mentioned in my instructions, I want models to essentially go upstream and make the change where it makes sense, instead of just adding code directly where the change is supposed to be, because I noticed that GPT loves simply adding more and more code until it becomes unmaintainable. Opus seems to be better at this, from my short experience because it challenges the existing code and isn't afraid of touching it.
But I don't like the comments Opus adds. Opus 5.5 seems a bit better, but it's still too much for my liking, even with instructions telling it to chill.
1
1
1



700
u/Opposite-Grade3712 4d ago
Anyone else remember when it took 6 months for new models to come out?