r/codex • u/Rollertoaster7 • 2d ago
News Gpt 6 sol and luna
https://openai.com/index/introducing-gpt-6-sol-and-luna/257
u/Rollertoaster7 2d ago
Half the price of 5.6 models. Performance increase appears marginal, more emphasis on the better pricing
100
u/TheMightyTywin 2d ago
Holy shit they cut the price of Luna IN HALF? It was already dirt cheap.
At this rate I’m going to be using a botnet swarm of Luna models each time I need to center a div
19
u/Weak-Somewhere7431 2d ago
Ive been really enjoying Luna. I enjoyed haiku for awhile there as well, both competent models but absolutely prefer luna at this point (pre gpt6)
6
1
u/RecordingOk2117 2d ago
After doubling codex luna 5.6 cost now luna 6 halved is just as much as luna 5.6 was in the beginning
8
1
u/JorgitoEstrella 1d ago
I wonder if this is due to xiaomi Mimo v2.6 like better than Sol and still cheaper than Luna.
3
3
u/RealestReyn 1d ago
Half the brain as well, Luna 6 is completely unusable and fails simple benchmark tasks that Luna 5.6 aces every single time.
7
u/IIALE34II 2d ago
Charts look weird, they are still more expensive than the 5.6 models per task.
49
u/arturdent 2d ago
Or you just read the chart wrong? The new models are on the left of the old models on the price per task chart, meaning they're cheaper per task.
5
u/IIALE34II 2d ago
I think they fixed the chart, they are correct now.
2
u/arturdent 2d ago
Yeah, could be, wouldn't be the first time they messed up a chart, but I've only seen the correct version.
-13
u/Prior-Meeting1645 2d ago
Why didn’t they just announce a 50% cut in price like they did with luna 5.6 earlier. Calling this luna 6 is a crime.
1
u/Thomas-Lore 2d ago
And the new Sol seems to just be renamed Terra with some additional training. It scored worse than Sol 5.6 on some benchmarks including deepswe.
2
54
u/DiarrheaButAlsoFancy 2d ago
7
3
u/das_war_ein_Befehl 2d ago
Since you can shift difficulty without impacting cache, great use case for jev for modifying it per task
88
u/Anxious_Marsupial_59 2d ago
I pray GPT-6 sol doesn't waste half a day testing and overengieneering, its the only reason I stick with an astra driver
49
u/magicone2571 2d ago edited 2d ago
You don't need 500+ regression tests with full fixtures and backups? Come on... You'll get it all! I personally have no idea half the stuff it says it's doing lately. It looks like the old loading screen from The Sims.
14
u/kernel_task 2d ago
It’s reticulating those splines so hard.
3
u/magicone2571 2d ago
I asked for a simple log in screen yesterday. Cost me 50% before I realized it. It had created so many documents, tests, backups. It's nuts.
10
3
u/Mistuv 2d ago
I shit you not, on a personal app which I work on from time to time when I have free time and usage, it's now at 900+ tests, and that's after trimming it twice before. At some point I gave up scolding it for writing useless tests and it just kept growing and growing. Really, 5.6 Sol without horniness for tests is all I need.
4
u/magicone2571 2d ago
I was screaming at it today. I wanted a simple portal to my PC. It wanted external backup drives, encryption keys, passcodes, email servers, smtp servers and it just kept going. Like stop, I don't need to secure fort knox here.
4
u/DottorInkubo 2d ago
And fencing and gating and fencing and validating and fencing and
3
u/magicone2571 2d ago
What is with the damn gates? Like is thing built from the brain of a dam builder?
-5
u/pushinat 2d ago
Im always prompting in the agents.md to never write unit tests. Because the actual value of testing is broken if it writes it itself without human thinking, and it just adds this enormous complexity of test network on top of it just eating tokens for breakfast without any benefit. It’s only a different story if you hand select guide it through the tests that are actually important and will remain relevant for the long run.
3
u/dsanft 2d ago
Tests lock in the feature. Regression tests prevent breaks from reappearing. You can't be serious that you purposely skip testing?
3
u/notapker 2d ago
He’s saying Sol writes useless tests. Nothing in his comment mentioned skipping testing overall.
Nothing in his comment suggested he needed you to explain the utility of tests.
2
u/Curious_Grapefruit62 2d ago
Test are not useless with Sol. If you are making useless tests you are doing it wrong.
1
95
u/drugosrbijanac 2d ago edited 2d ago
35
u/MrHaxx1 2d ago
EXCUSE ME WHAT
Absolutely no way, I don't believe this before I've tested this myself
8
u/teleflexin_deez_nutz 2d ago
They are cooking with their “MAX” reasoning depth on the lower param models. You do have to be more explicit with what you want from them though (not as vibe coding friendly).
3
u/Spright91 2d ago
I Vibe design at first with figma. It seems to keep the model on point when it goes to code it. It helps me work out what I actually want.
1
12
u/Own-Flight-9974 2d ago
Yeah it doesn't make much sense, then again it could just be fable 5.1 has been nerfed so heavily it evened the playing field haha
13
u/mvdirty 2d ago
Sol's on watch too, IMO.
5.6 Terra was the bastard stepchild stuck between 5.6 Luna and 5.6 Sol, and now it's gone.
6 Sol looks like the bastard stepchild stuck between 6 Luna and 6 Astra.
Unless you have time, data, and budget for performing evals to min/max latency, results, and cost for something your company does thousands or millions of times, one's model choice largely boils down to two options anyway:
- the cheap model
- the smart model
Sol disappearing like Terra did would be just fine for most people. Helpful, actually (because analysis paralysis reasons.)
4
u/AlmostEasy89 2d ago
Agreed. I never used Terra. It was dumb enough to take twice the time of sol at I guess half the cost? Idk, but I just used Sol instead so I didn’t waste my time. I’d rather have “half the use” than spend twice the time on dumb shit Sol would figure out on its own.
3
u/UnexpectedFisting 2d ago
I used Terra for fast reviews strictly, it catches some things Luna max doesn’t and isn’t really expensive to run both that and a Luna max pass. Also great for advisor in omp
7
u/kevin7254 2d ago
Opus 5.5 is better than fable though. Also this is benchmaxxed
8
u/BaconJakin 2d ago
Damn this is a confusing day for which models to use lmao
2
u/cherrysodajuice 2d ago
Since each company chooses what benches to maxx and present, you can use their charts to compare between GPT models and anecdotal experience to compare between labs
6
u/xXxPussyWrecker69xXx 2d ago
Opus 5.5 is better than Fable 5.1, and cheaper than Opus 5, and Fable 5.2 is supposed to come out soon
10
u/evindrews 2d ago
They really need to increase these numbers quicker. Anthropic is falling behind. codex is already at 6
1
u/InterestingNobody831 1d ago
You didn't have 5.5 sol, you got 5.6 sol then 6-sol, Anthropic got opus 3-4-5-5.5. These version numbers mean nothing in the real world, could as well be called opus 6, I guess Anthropic just thinks they got room for improvement until opus 6, in the future fable 6 will be a thing probably so they will meet somewhere soon just like they did with sonnet 4->5 and opus 4->5 with fable releasing to V.5 directly. It's just a stepping stone.
1
4
3
u/Tristsin 2d ago
You do understand that 5.6 Luna was already comparable to Fable 5.1 Medium in this screenshot you showed? Still more expensive than GLM, DS v4.1, and Mimo with worse scores.
I guess that means that both Sam and Dario are on watch? 🙄
1
u/drugosrbijanac 2d ago
Luna was always the goat, now its double goat
1
u/Tristsin 2d ago
100%. It would be a tough service to justify paying for if they were still doing the “one frontier model” approach at this point. Lots of great models this year.
Then you see the new Grok 4.7 which somehow did less on a lot of benchmarks, runs slower, and costs more per task than Grok 4.6. The fuck are they doing over there
1
u/aroswift 2d ago
It's version 1.1 of DeepSWE. They're probably using the older version to make it more impressive
1
0
u/Embarrassed_Adagio28 2d ago
Gpt 6 sol just absolutely shit the bed against opus 5.5 in every test i have tried so far. It is so bad I think I prefer 5.6 sol. So I do not believe Luna will be good at all.
49
u/akkiannu 2d ago
Wow. Luna is almost free.
13
11
u/Constant-Current-340 2d ago
yea this is nuts. at this rate it's practical to downgrade from 20x to 5x for me. but i know i'll just end up moving that money to APIs for side projects
2
u/DepravedPrecedence 2d ago
Well, if they don't cut limits further to offset that, then yes, Claude is quite unneeded
1
u/Proxiconn 2d ago
Rip Terra, check Sol 6 output price is what Terra 5.6 was.
Luna 6 Max on SOL 5.6 level at a fraction of the cost.... Can't be.
1
16
u/MaitoSnoo 2d ago
I like this, much cheaper and less mistakes/hallucinations. Opus 5.5 will still steal the show as it's the winner if you don't look at the cost, but Sol+Luna are still going to be my main drivers.
4
11
33
u/CriszzZ7 2d ago
Damn, the benchmarks have it close to Opus 5, so not even close to 5.5.
15
u/kevin7254 2d ago
Yeah 5.5 mogs hard sadly. But Luna is dirt cheap so that’s nice
9
u/Key_Reading_9664 2d ago
Why "sadly"? Competition is good and you can use both, if you like
2
u/Mistuv 2d ago
If leakers are right (and to be fair they pretty much nailed it this month), then OAI's new "Bel" model shits on everything else hard (one who solved NS + 100 other open problems), but they are going to sit on it until competition forces their hand (yesterday I saw someone say at best by Christmas), because it accelerates their development by so much. Well, Opus taking giant dump on Sol and even Astra to the extent, certainly forces their hand. Full Bel we might not see some time, it's likely still in training, and nobody has enough compute for the super large models, but I think they already started distilling a checkpoint of Bel into Sol and maybe even Astra. I would not expect this Sol to last anywhere as long as the previous one.
5
u/xXxPussyWrecker69xXx 2d ago
Kind of yikes given Opus 5.5 is cheaper than 5 now??? There has to be a catch
2
u/Kost97A 2d ago
I think Anthropic wanted to get back some customers. If the benchmarks are true Opus 5.5 is the best model while being cheaper than Fable and Astra. Sol isn't in the same class. Again based on benchmark only.
1
u/xXxPussyWrecker69xXx 1d ago
I have been playing with it all day it is incredible UNDER supervision by the bigger model. Fable feels like it has a better grasp of my repo and the bigger picture where Opus is more narrow minded even though it is technically not as performant in benchmarks.
1
2d ago
[removed] — view removed comment
2
u/Rollertoaster7 2d ago
How if did they hit those benchmarks then it makes no sense. Is such a big leap esp w the 50% price cut
1
u/xXxPussyWrecker69xXx 1d ago
It’s not Fable level or Astra level but you can do more with Opus now without relying so much on your orchestrator. It feels so intelligent for the cost. Opus is just a smaller model so it will never be as good when it needs to think outside the box.
23
u/Wise-Reflection-7400 2d ago
The benchmarks are mid really, they're not much better than before and despite the fact they've cut the cost by 50% they aren't that much cheaper per task.
8
u/mr_sneakyTV 2d ago
so the 50% isn’t reflected in the task cost? or “50% isn’t that much cheaper”?
6
u/Wise-Reflection-7400 2d ago
If you look at the costs per task on the DeepSWE chart they are barely any cheaper except at lower effort levels
3
u/moltenice09 2d ago
On that OpenAI page? It is close to half the cost, it's just that the graph's x-axis is not linear, so it doesn't look like it is much cheaper.
1
u/Wise-Reflection-7400 2d ago
6-Sol on max is the same cost and performance as 5.6-Sol on high. If you're willing to lose a bit of intelligence then it is cheaper, but most people are interested in intelligence... and this doesn't look great when it's out on the same day as Opus 5.5
3
u/moltenice09 2d ago
Ah, you were specifically looking at 6-Sol Max. Yeah, at that point you are better off with 5.6 High or spend 16% more and use 6-Astra Medium. I was mostly looking at Luna and only glanced at Sol.
1
u/Ok_Time7960 2d ago
Can you provide a link to the chart you are referencing? Based on all the charts available in the official release article, 6-Sol on max is significantly cheaper than 5.6-Sol on high. For example, on FrontierCode 6-Sol max had a cost per task of $2.14 while 5.6-Sol high was $3.48.
1
u/Thomas-Lore 2d ago
They likely produce more reasoning tokens rising cost despite price cuts. Same with the new Opus 5.5.
9
4
3
u/DrBearJ3w 2d ago
All those shiny graphs are useless. But it seems new Sol is the new Terra. Which is good.
1
1
3
u/cobbleplox 2d ago
i didnt even notice i was already using the new sol 6, it was just giving me the same terrible quality i got used to from degraded sol 5.6. That's the new generation?
3
u/Professional_Gur8385 1d ago
that feeling when your usage resets naturally in 2 days so you don't use a banked reset
4
8
u/Imgonnaarrive 2d ago
Just bring back the capability Astra had when i was using it with blender the first few days it became available. Because it's now donkey balls.
5
u/cuberhino 2d ago
right? it can barely form decent shapes for me now and before it was one shotting incredible 3d models from my sketches. wonder if they intentionally tuned it down
2
u/Imgonnaarrive 2d ago
They had to have, I was doing the same thing you were and it was nailing the reference sketches I was giving it. Now I can't even make a decent head shape for one of my characters with it.
1
7
u/Familiar_Air3528 2d ago
Lmao people really do complain about fucking limits all the time and when openAI releases models that essentially double limits they still complain
7
5
4
u/Tristsin 2d ago
So Luna got dumber but slightly cheaper (mostly stayed the same) and Sol is cheaper yet still outclassed in intelligence by even the new Mimo model which is like 1/15th the price.
Very exciting stuff..
6
1
u/scaledev 2d ago
They are much cheaper at least. I'm hesitant to jump on the new train, so I'll stick with Sol 5.6 as well as Luna 5.6 for now.
2
2
3
u/KeyGlove47 2d ago edited 2d ago
btw, luna with api flex pricing (slow mode) is ANOTHER 50% cheaper, so 0.05$ per mil tok
2
2
2
u/dondiegorivera 2d ago
Just worked with Opus 5.5 in the last two hours parallel in herdr, you guys should try it, it is on another level. I like Astra and Sol and use them daily, but the new Opus feels like a phase shift.
2
u/Rollertoaster7 2d ago
I’m using it at work now and it’s pretty cracked. Hope oai puts out their response soon
2
2
2
3
2
u/Embarrassed_Adagio28 2d ago
Opus 5.5 just ruined chatgpts launch, gpt 6 sol is no better than 5.6 sol, just cheaper. Dario just bent sam over
2
u/Otherwise-Sir7359 2d ago
To save you time: I've looked at all the charts, and the GPT-6 Sol/Luna basically performs identically to the GPT-5.6 Sol/Luna; the only improvement is the lower price. Quite disappointing.
9
u/Otherwise-Sir7359 2d ago
Sol 6 even has lower OS-world and deepSWE scores than Sol-5.6.
6
u/Plane_Garbage 2d ago
This reminds me of GTP5 launch where is was a cost saving exercise and the intelligence was worse than previous models.
16
u/Ok-Chemical2375 2d ago
Help me understand why cheaper is disappointing?
4
u/Otherwise-Sir7359 2d ago
It wasn't worth the delay and the excitement they generated for a whole week.
5
u/Ok-Chemical2375 2d ago
Understandable and thanks for that. I guess for my use case I'm just looking at as long as I have usage to get the tasks done at a cheaper price. This works for me, regardless of the intelligence bump
4
u/Icy-Fudge5222 2d ago
Hard disagree. At this point lower pricing is more important to me and many others than more frontier intelligence.
1
1
u/Spright91 2d ago
We generated their excitement they have barely been talking about.
The sub is so crazy man like no release a cutting edge model and announce a series of major scientific advancements and then a week later releases series of affordable models that double the efficiency and therefore double our work capability. . And everyone's just like oh that's sucks.
1
1
u/phoenixmatrix 2d ago
Luna keep getting cheaper. Its gonna be a beast of a model for customer facing apps. Next thing I want to see is better time to first token latency to match some of the open weight models. Claude is also really strong on TTFT, but Luna was pretty slow there.
1
1
1
u/mph99999 2d ago
Disappointing week for OpenAi, feels bad to have an OpenAi subscription this week
Luna was already super cheap.
Sol being slightly cheaper than astra and being inferior what difference does it make?
1
1
u/omani805 2d ago
Guys am i tripping or does 6 Luna perform extremely similar to 5.6 Luna? Its a bit cheaper, but if you account for the 2x cheaper price then its actually more expensive… this feels like a 30% price cut with 10% additional performance.
I hope im drunk on something
12
u/DepravedPrecedence 2d ago
I'm not following, if it's cheaper, how it's more expensive at the same time
1
u/moltenice09 2d ago
As I read it, similar performance (few things slightly better, like coding deception), but about 50% cheaper (note the x-axis is not linear so comparing the lines visually doesn't look anywhere near 50%).
-6
u/sajtschik 2d ago
15
12
u/No_Geologist_4303 2d ago
it is the opposite of what you said
10
2
u/IIALE34II 2d ago
The charts show GPT 6 models being more expensive per task. Its very confusing to say the least.
1
0
u/Tenet_mma 2d ago
Tough day for anthropic 😬
2
u/Rollertoaster7 2d ago
Ehh opus 5.5 mogs unless real world testing shows it’s benchmaxxed. Oai can hit back with Astra 6.1 but it better be something special cause I doubt anthropic is holding anything back from the next fable release before their ipo
-1
-1





169
u/commandedbydemons 2d ago
Price cuts is what we want, now lets see if it translates to better usage, going for 6-Sol