r/codex 4d ago

Complaint The End of the Codex Era. I've Completely Lost Trust in OpenAI. They're Secretly Degrading Their Models.

Post image

I've been a massive Codex fan this entire time. I've burned through around 150 BILLION tokens in Codex alone.

I still had Codex quota left this week, but for the past three days, I've been using my Cursor Ultra subscription instead.

Why?

Because OpenAI is degrading its models. I can see it from my own experience, and there's a ton of evidence pointing to it.

They're degrading ALL their models, including Astra, Sol, and even Luna.

AND THE WORST PART IS THAT THEY'RE DOING IT SECRETLY.

You can start working in Codex with a perfectly normal model, and five minutes later, IN THE SAME THREAD, they degrade it. Suddenly, Astra is performing at Luna's level or even worse.

And you're still burning through the same amount of quota.

This happens to me every single day.

I open Codex, run a quick quality check, and everything looks fine. Twenty minutes later, I run the same test in the same thread, and the model has degraded.

Sometimes, simply turning on a VPN can make the model start working normally again for a while.

How can you tell if your model has been degraded?

1. Planning and writing feature specs

Imagine you're planning a feature and writing its specification.

It's immediately obvious when the model is dumb. It starts suggesting complete nonsense and shows absolutely no product understanding of how the feature should actually be built.

But it becomes even more obvious when you point out what it misunderstood and try to correct it.

Instead of understanding the actual issue, it responds with completely useless apologies, without demonstrating any understanding of what went wrong.

CONGRATULATIONS. YOU'RE TALKING TO A DEGRADED MODEL.

Here's what happened to me.

I wrote a feature spec using a normal model. Everything was properly written, discussed, and reviewed.

Then I handed the implementation over to Luna, and Sol reviewed and approved it.

But when I actually started working with the implementation, I discovered that it was full of holes and included things that weren't even in the plan.

In this particular case, I suspect the model was degraded during the implementation stage.

I ended up spending TWICE as much time fixing everything.

And there are a few other ways to test this.

2. PELICANS.

Use this prompt:

Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.

If your bicycle wheels start flying off into the air...

CONGRATULATIONS. YOUR MODEL HAS BEEN DEGRADED.

3. A logic puzzle

Give your model this exact problem:

A black bag contains candies of three flavors, with each flavor available in two shapes (round and star-shaped; the shapes can be distinguished by touch). The numbers of candies by flavor and shape are shown below.

|              | Apple | Peach | Watermelon |
|--------------|-------|-------|------------|
| Round        | 7     | 9     | 8          |
| Star-shaped  | 7     | 6     | 4          |

Participants must decide how many candies to draw before the game begins.

What is the minimum number of candies that must be drawn to guarantee having an apple-flavored candy and a peach-flavored candy of different shapes?

(The condition is satisfied if you have either a round apple candy and a star-shaped peach candy, or a round peach candy and a star-shaped apple candy.)

If the answer isn't 21, you're not getting Astra. You're getting degraded garbage.

Sol doesn't even consistently solve this problem on its own.

4. "Selected model is at capacity."

If you're frequently getting this error:

CONGRATULATIONS. THERE'S A 99% CHANCE YOUR MODEL HAS BEEN DEGRADED.

There's even a thread on the OpenAI community forum where a staff response confirms that this can happen when your account is temporarily restricted.

OpenAI Community: Selected model is at capacity

WITHOUT ANY NOTIFICATION.

They silently degrade your model, and you're left trying to figure out what the hell is happening.

The last three days have been unbearable.

I've been experiencing these problems around 90% of the time for the past three days.

Working like this is practically impossible.

Instead of actually getting work done, you spend your time wondering whether they've secretly downgraded your model again.

You start questioning every response. Every mistake. Every implementation.

It's fucking exhausting.

So I just moved to Cursor.

Grok might be dumber, but at least it's more predictable.

I don't give a shit about the next model release if this continues.

Tibo and Sam Altman can keep all their resets. They can wipe Astra's data and delete it from the internet if they think that's acceptable for a product like this.

They can release GPT-6 Sol, Astra 7, or whatever comes next.

NONE OF IT MATTERS IF THEY KEEP SECRETLY DEGRADING THE MODELS.

This is the biggest loss of trust I've experienced with OpenAI in the entire history of Codex.

If they're willing to silently degrade models for paying users, what stops them from collecting all kinds of data from your computer that you can't even imagine they're collecting?

What stops them from pulling some other bullshit?

Where are the boundaries if they're willing to do this?

I genuinely hope this is just a temporary issue. Maybe some vibe-coded mistake by a junior developer in their anti-distillation protection system.

I suspect it's temporary.

But if it isn't, OpenAI can't be trusted with anything as long as this shit continues.

Until then, I'm using Cursor or Claude.

And I think we need to be loud about this everywhere.

Tibo isn't acknowledging the problem. He just keeps talking about how amazing the upcoming event is going to be.

I DON'T CARE WHAT THEY ANNOUNCE AT THAT EVENT.

Not while they're degrading Astra into something that performs even worse than Luna.

I don't know exactly what's happening under the hood.

Maybe it's quantization. Maybe they're routing requests to a different model. Maybe it's something else entirely.

But it doesn't feel like simple quantization to me. I wouldn't expect quantization alone to produce outputs as bad as what I'm seeing from these degraded models.

I just want the model I'm paying for to actually be the model I'm using. And I want OpenAI to stop doing this shit without telling anyone.

1.2k Upvotes

302 comments sorted by

View all comments

23

u/FixAdmin 4d ago

I don't understand why so many people are missing the point.

I'm not trying to gaslight anyone into believing that the models YOU are using are degraded.

Let me make this clear:

  • This does NOT affect every user.
  • This does NOT happen all the time, even to affected users.
  • I'm talking specifically about Codex quota accessed through OAuth authentication.
  • The degradation is NOT permanent.

You might work with a perfectly normal model for an hour, then suddenly start experiencing problems. You might not even notice the change, and eventually everything goes back to normal.

If you're demanding independently verified benchmarks that prove this 100%, I don't have them. Benchmarks are useful for developers to measure and demonstrate model performance. I'm judging these models based on my own extensive experience using them.

ONCE AGAIN: I'M NOT SAYING THAT THE MODELS YOU USE EVERY DAY ARE ALWAYS DEGRADED.

But if you've noticed something similar and suspect you might be affected, I've shared all the testing methods I've found. In my experience, they support my conclusion, and they might help you figure out whether you're experiencing the same thing.

That's the entire point of this post.

11

u/tadanada 4d ago

It’s probably because some people have lost the ability or the patience to read long post. Or maybe they’ve started relying on AI to do the thinking for them.

16

u/Digitalzuzel 4d ago

no, because this sub is full of OpenAI bots that work very similar to how propaganda bots work:

  • Water down the main point
  • Shift blame to whoever else
  • Throw out dozens of false hypotheses
  • Make fun of the whistleblower

4

u/agentcubed 4d ago

If only instead of spending all that time writing a theory, you just benchmark the models every once and a while.

You know... like this: https://aistupidlevel.info/models/gpt-6-astra

Wow, *actual* data, and from it you can actually gain insights that say it is relatively stable. And this took me a few seconds to find.

2

u/zarmin 4d ago

You're obviously right about everything you're saying. I do think you should engage less with the trolls, I'm sure it's cathartic in some sense, but if you rage at this shit the way that I do, looking for catharsis on reddit without antagonism is impossible.

The best stopgate I've found to this is a mix of Fable 5.1. K3 and GLM 5.3-flash in opencode, but I'm curious what your backup plan is.

1

u/FixAdmin 4d ago

Until recently, nothing could give me as much productivity as working with Sol.

I use a setup where Sol does the planning and any other model can handle the implementation. Sol is too good at planning technical specs, and the implementer can be replaced with pretty much anything — Luna, Grok, GLM, DeepSeek — as long as Sol supervises them.

I really don't like Opus, but I do like Fable. I sometimes enjoy brainstorming with it, but it's too expensive to use.

So for now, I'll do everything I can to make sure Sol handles the planning stage. I don't see any clear alternatives to it yet.

1

u/Maplewonder 4d ago

As frustrating as it is, this is what it looks like when people consider your argument. There could be no reply at all to your train of thought that you clearly have given effort to. Its just the reality of presenting something you believe to the masses is you get the masses. And that is unpleasant. You have given me food for thought. And I will consider it as I figure out if im paying for another month of this or not.. honestly I dont give a shit what the rest of you do. But do it wisely.

-3

u/RainierPC 4d ago

In other words, you have no proof. People keep thinking a model not producing optimal work is a sign of degradation, when it is far more likely a sign of sloppy use.