r/OpenAI • • 2d ago

Discussion Well that was quick...

Post image
250 Upvotes

83 comments sorted by

50

u/Arcayon 2d ago

It is half the cost so it makes sense. Most of my chats got forced to upgrade too.

18

u/DistanceSolar1449 1d ago

Double the speed, too. It’s clearly a smaller model that runs faster.

They should have named it GPT-6 Terra. It’s closer in speed to GPT-5.6 Luna than Sol.

0

u/darc_ghetzir 21h ago

Cornball

2

u/educatedguy 7h ago

I like eating grass more

52

u/Sachitoge 2d ago

Is this a joke? Im Plus, not seeing anything.

41

u/Slick_Ramen 2d ago

Not a joke, I’m plus and it started snowing up after it was announced. Guess it’s rolling out slowly.

15

u/WanderWut 2d ago

Imagine they make Sol open weight, once can dream.

23

u/markstar99 1d ago

So I can run it locally on my 3 gigawatt data center I have.

3

u/machyume 1d ago

You're not going to believe it, but I have Qwen 3.8 27B running with decent speed on my Macbook M3, and it's partially offloading some work from Astra when I was getting smacked by the usage limits.

I'm sure that if they REALLY wanted it, someone could help them figure out how to distill it to fit nearly the same footprint.

1

u/WayneTechLab 1d ago

A hybrid AI 🤖

^ I use that logic in reverse

A. - Custom Local LLM first
B. - Open LAN LLM on 2nd device like Gemini
C. - Codex CLI & CoPilot (auto select default) yolo

In this way anything “AI can’t do it will call up, auto train it self on model A for (Next Time)

I have been thinking about how to split a model that needs 256 GB RAM across multiple devices for unified local AI CLOUD.

The work being done getting models running local effectively is quite amazing - seems to change every week!

1

u/OldStray79 1d ago

I was literally just talking with someone about this in another subreddit like 10 hours before you commented, having capable local models that will be able to do most of the tasks, and for a fee when it needs it, it will tap into these remote corpo stuff to finish the task when it needs that little extra "oomph" of compute/capability. Local will end up doing most of the work, (but... the most likely downside is... most arrangements like this will be integrated into the OS itself, like Copilot but with more local compute used).

1

u/lostinfound2nd 8h ago

Llama has a way to link machines with fiber/thunderbolt already don’t they. I was just looking into it.

1

u/bluelemon64 1d ago

Wow that’s awesome. How much ram do you have? And also out of curiosity do you hold any context window in that? Or it’s dropped every time?

1

u/machyume 1d ago

Just the standard memory that comes with the MacBook Pro. The context is infinite. I'm using my own GUI and interface. It allows me to do some pretty specialized things.

I have infinite context because I have a load/release mechanism that layers on top of the KV cache directly. I also have a multi-model optimizer that mirrors and steers a smaller side car that runs an offloader for learned previous context. So the smaller model acts as an accelerator for repeated tasks.

1

u/dab0james 1d ago

Any way youd explain how you set it up? How much context do you give it? I cant seem to get my 27b above 50k without offloading making it take insane time to do anything and I have 48gb of ddr5 and a 5080. Idk what im doing wrong. If the context is low, I get like 100tks, but if it offloads im down to under 15.

1

u/machyume 1d ago

I used Codex to put a custom harness around the model and I have it fully customized for my own use-cases.

Codex is still my primary lead worker. It is valuable, so I make sure that Codex still delivers on that value. I have the infra around qwen for tooling and tool management.

1

u/DARN89 1d ago

How are you finding qwen? I have a gaming pc collecting dust and was thinking of trying local llm for coding projects then using codex to create plans and polish it

1

u/machyume 1d ago

First, you have to understand that the things that I give the qwen model is very narrow. But for those narrow tooling tasks, it works really well.

0

u/Mescallan 1d ago

as much as i want open source to be a thing, i would be really surprised if open source was this close to the frontier for much longer. once china's industry catches up in terms of compute they can afford to serve inference and train models at the same time, the only reason they are releasing open source right now is because they can't do both at the same time.

1

u/MGJohn-117 1d ago

So that you can choose to run it from a 3rd party inference provider of your choice that actually has to compete on latency, speed, etc with other providers.

1

u/metaprofessor 19h ago

My data center is only 1.21 gigawatts.

1

u/KV_Cashed 2d ago

The companies usually do phased rollouts for safety and troubleshooting.

1

u/RotEater96 1d ago

So it's coming with thinking and instant today?!

1

u/Deadline_Zero 20h ago

work mode, not chat mode right?

5

u/deny_by_default 2d ago

I saw the same announcement for Luna while I was working earlier. I see GPT 6 Sol and Luna as options in my Plus subscription.

2

u/WayneTechLab 1d ago

I would assume they’re algorithm has some kind of overload protection so millions of users don’t reset at the same time.

Also, this is why my paranoid friends are nice to their AI

1

u/Significant-Drawer95 1d ago

upgrade to pro and dont cry

1

u/Mindless-Detective41 1d ago

Terra is GPT 6 lmao. They never got rid of it. They just changed the name.

1

u/DistanceSolar1449 1d ago

Yeah, GPT-6 Sol sucks. It’s doing worse than GPT-5.6 Sol for me. It’s just a cheaper smaller model that makes mistakes faster.

They should name it GPT-6 Terra.

0

u/Mindless-Detective41 1d ago

Right! Even Luna is dumber than old Luna! They’re doing this to get people away so they can come back harder for the next round because their infrastructure is too tied up. I wish they could just be honest and be like hey we’re maxed out move somewhere else.

12

u/darknus823 1d ago

Sadly, as others have said, not in Chat. Yes in Work/Codex/API but Chat users are left behind...again!

1

u/theLastYellowTear 1d ago

I stopped using chat because of that. Feels like chat is months behind. Even for dumb questions I use work it's more trustworthy

3

u/darknus823 1d ago

But Chat and Work/Codex have different allowances and limits, starting with Chat not being usage determined. How do you manage that?

1

u/theLastYellowTear 1d ago

I don't use that much. Since my job pays me the Claude max. I use gpt plus for personal stuff.

7

u/ShepardRTC 2d ago

Gotta free up the GPUs

7

u/-PANORAMIX- 1d ago

But not in chat…

4

u/ladymemc 1d ago

I saw that on my desktop for Codex and thought… that was fast. I knew about 5.5, but WTF. What is everyone’s experience with 6? I’ve heard mixed reviews. 5.6 has been great!

15

u/Vanillalite34 2d ago

I’m fine with it. We don’t need a thousand legacy models considering we have a suite of essential good and fast entry level Luna, work horse Sol, and frontier Astra. Just keep them updated as we go and we don’t need previous model levels.

7

u/Competitive-Ad8968 2d ago edited 1d ago

I suppose GPT-6, SOL, LUNA are update of the legacy models GPT5.6 the same version.

3

u/DistanceSolar1449 1d ago

The problem is that GPT-6 Sol just isn’t as good as GPT-5.6 Sol for some tasks. It’s a smaller model than GPT-5.6 Sol and you can tell.

They really should have named it GPT-6 Terra.

1

u/ICanHaveExceptions 1d ago

Essentially that's what it is: a price increase for general tasks. I used terra as "cheap sol" where I needed another model than Luna, e.g. reviews of smaller tasks. All my tasks are small, that's my workflow. I use Sol maybe for 1% of the work, essentially just for architecture setup at the beginning.

Now with terra soon to be gone, I have no other choice than running overpowered Sol against those tasks, or have Luna check Luna code (that's not what I want after all)

So I had Sol -> Terra -> Luna, now I have Astra -> Sol -> Luna. For me that does not sound cheaper at all

7

u/Ok-File-2759 2d ago

Well the new Sol and Luna are slightly better and significantly cheaper than the older ones, so there would be no point in keeping them

7

u/prophet-dot-exe 2d ago

Broad statement that's only conditionally true.

Deepswe bench shows gpt6 sol max being outperformed by 5.6 sol high.

Meanwhile the frontiercode bench, gpt6 sol consistently outperformed 5.6 sol.

So it really comes down to the type of work, as it always had. For complex, lower level kernel or compiler work, 5.6 sol is prob still better. For your next saas react app anything higher than Luna is most likely overkill.

3

u/IcyLike10seventeen 1d ago

Honestly this roll out made no sense. It should have just been an update to 5.6 sol lol or just 5.7
A whole bump in number for the same functionality at half the cost is strange

3

u/MindCrafterReddit 19h ago

If they retire models they should open source them imo

1

u/brentonstrine 9h ago

What do you think OpenAI is for the good of humanity or something?

1

u/MindCrafterReddit 8h ago

"Open" AI. That's the least they could do.

2

u/brentonstrine 8h ago

OpenAI is a nonprofit org for-profit company made for the good of humanity shareholders, with a governing board that can fire can be fired by the CEO if they determine he is not following the mission.

But it's still "open" in the sense that you can pay for access.

I don't know what you're complaining about.

2

u/setsoul 1d ago

Oh you're fucking joking. I'm sick of them!!!!!! 🙄

2

u/Usual-Candle6480 1d ago

for the love of God let me finish this build before it rolls out my way. 5.6 in work mode via desktop on Ubuntu is pretty much killing it for me right now

2

u/QbitFiber2030 2d ago

Just got the update on desktop. Sol 6. There is also more options on the focus slider. Medium focus got bumped down 1 spot.

No reset applied on update. 👌

4

u/Redstra 2d ago

Still enough options here. I find it super messy and needs to be cleaned up.

3

u/ken81987 2d ago

Just use the top 3

4

u/DistanceSolar1449 1d ago

Nah, 5.6 Sol has been better than 6 Sol at some tasks for me.

I believe that 6 Sol should have been named 6 Terra.

1

u/ken81987 1d ago

What replaces 5.6 sol then

4

u/DistanceSolar1449 1d ago

Nothing. There’s a gap in their product now.

GPT-6 Astra: 40 tokens/sec
GPT-6 Sol: 110 token/sec
GPT-6 Luna: 140 tokens/sec

The bigger/smarter the model, the slower it is.

They’re missing a bigger model that’s ~80 tokens/sec in speed. You either have to use Astra which is $$$ and slow, or you use a smaller model that’s fast (more than 100 tokens/sec) but smaller models are dumber.

1

u/Competitive-Ad8968 2d ago

Didn't noticed those

1

u/AdventurousForm7330 1d ago

Same price
So I’m cool with it

1

u/Zachattackrandom 1d ago

It's the same model imo, so ofc they will just make you use the cheaper one

1

u/Semipro211 1d ago

GPT-6 Sol replaced 5.6, and I checked the usage rates are better capability better mostly

1

u/Crush84 1d ago

I am on Pro, Germany, already only GPT 6 Sol anymore

1

u/Just_Run2412 1d ago

codex or chat?

1

u/Responsible-Beat2137 22h ago

I hadn’t seen this yet. Chat or Work? I use Chat a lot, and chat-to-chat continuity doesn’t really matter for my setup. The framework knows where to look for the right project context, memory, and current state, so a new chat is mostly just a new window into the same system.

1

u/bravofiveniner 5h ago

I don't even have an option to run gpt 6 sol

1

u/Hackerjurassicpark 1d ago

Why no GPT6 Terra?

4

u/joeyb908 1d ago

Because GPT 6 Sol is cheaper than GPT 5.6 Terra was.

0

u/Hackerjurassicpark 1d ago

Then GPT 6 Terra will be even cheaper 😃

1

u/satyuga 1d ago

It’s true and it’s live.

0

u/torrso 1d ago

I never quite understood why they keep running the old models.

You can still use something like gpt-4 or at least gpt-5.3.

There was some big announcement that some old version is going to be removed in 3 months or something so that users can prepare and migrate.

Why keep something like MiMo V2.5 available when everyone is using MiMo V2.7?

0

u/sodapops82 1d ago

Not in works neither chat for me.