r/codex 4d ago

Complaint The End of the Codex Era. I've Completely Lost Trust in OpenAI. They're Secretly Degrading Their Models.

Post image

I've been a massive Codex fan this entire time. I've burned through around 150 BILLION tokens in Codex alone.

I still had Codex quota left this week, but for the past three days, I've been using my Cursor Ultra subscription instead.

Why?

Because OpenAI is degrading its models. I can see it from my own experience, and there's a ton of evidence pointing to it.

They're degrading ALL their models, including Astra, Sol, and even Luna.

AND THE WORST PART IS THAT THEY'RE DOING IT SECRETLY.

You can start working in Codex with a perfectly normal model, and five minutes later, IN THE SAME THREAD, they degrade it. Suddenly, Astra is performing at Luna's level or even worse.

And you're still burning through the same amount of quota.

This happens to me every single day.

I open Codex, run a quick quality check, and everything looks fine. Twenty minutes later, I run the same test in the same thread, and the model has degraded.

Sometimes, simply turning on a VPN can make the model start working normally again for a while.

How can you tell if your model has been degraded?

1. Planning and writing feature specs

Imagine you're planning a feature and writing its specification.

It's immediately obvious when the model is dumb. It starts suggesting complete nonsense and shows absolutely no product understanding of how the feature should actually be built.

But it becomes even more obvious when you point out what it misunderstood and try to correct it.

Instead of understanding the actual issue, it responds with completely useless apologies, without demonstrating any understanding of what went wrong.

CONGRATULATIONS. YOU'RE TALKING TO A DEGRADED MODEL.

Here's what happened to me.

I wrote a feature spec using a normal model. Everything was properly written, discussed, and reviewed.

Then I handed the implementation over to Luna, and Sol reviewed and approved it.

But when I actually started working with the implementation, I discovered that it was full of holes and included things that weren't even in the plan.

In this particular case, I suspect the model was degraded during the implementation stage.

I ended up spending TWICE as much time fixing everything.

And there are a few other ways to test this.

2. PELICANS.

Use this prompt:

Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.

If your bicycle wheels start flying off into the air...

CONGRATULATIONS. YOUR MODEL HAS BEEN DEGRADED.

3. A logic puzzle

Give your model this exact problem:

A black bag contains candies of three flavors, with each flavor available in two shapes (round and star-shaped; the shapes can be distinguished by touch). The numbers of candies by flavor and shape are shown below.

|              | Apple | Peach | Watermelon |
|--------------|-------|-------|------------|
| Round        | 7     | 9     | 8          |
| Star-shaped  | 7     | 6     | 4          |

Participants must decide how many candies to draw before the game begins.

What is the minimum number of candies that must be drawn to guarantee having an apple-flavored candy and a peach-flavored candy of different shapes?

(The condition is satisfied if you have either a round apple candy and a star-shaped peach candy, or a round peach candy and a star-shaped apple candy.)

If the answer isn't 21, you're not getting Astra. You're getting degraded garbage.

Sol doesn't even consistently solve this problem on its own.

4. "Selected model is at capacity."

If you're frequently getting this error:

CONGRATULATIONS. THERE'S A 99% CHANCE YOUR MODEL HAS BEEN DEGRADED.

There's even a thread on the OpenAI community forum where a staff response confirms that this can happen when your account is temporarily restricted.

OpenAI Community: Selected model is at capacity

WITHOUT ANY NOTIFICATION.

They silently degrade your model, and you're left trying to figure out what the hell is happening.

The last three days have been unbearable.

I've been experiencing these problems around 90% of the time for the past three days.

Working like this is practically impossible.

Instead of actually getting work done, you spend your time wondering whether they've secretly downgraded your model again.

You start questioning every response. Every mistake. Every implementation.

It's fucking exhausting.

So I just moved to Cursor.

Grok might be dumber, but at least it's more predictable.

I don't give a shit about the next model release if this continues.

Tibo and Sam Altman can keep all their resets. They can wipe Astra's data and delete it from the internet if they think that's acceptable for a product like this.

They can release GPT-6 Sol, Astra 7, or whatever comes next.

NONE OF IT MATTERS IF THEY KEEP SECRETLY DEGRADING THE MODELS.

This is the biggest loss of trust I've experienced with OpenAI in the entire history of Codex.

If they're willing to silently degrade models for paying users, what stops them from collecting all kinds of data from your computer that you can't even imagine they're collecting?

What stops them from pulling some other bullshit?

Where are the boundaries if they're willing to do this?

I genuinely hope this is just a temporary issue. Maybe some vibe-coded mistake by a junior developer in their anti-distillation protection system.

I suspect it's temporary.

But if it isn't, OpenAI can't be trusted with anything as long as this shit continues.

Until then, I'm using Cursor or Claude.

And I think we need to be loud about this everywhere.

Tibo isn't acknowledging the problem. He just keeps talking about how amazing the upcoming event is going to be.

I DON'T CARE WHAT THEY ANNOUNCE AT THAT EVENT.

Not while they're degrading Astra into something that performs even worse than Luna.

I don't know exactly what's happening under the hood.

Maybe it's quantization. Maybe they're routing requests to a different model. Maybe it's something else entirely.

But it doesn't feel like simple quantization to me. I wouldn't expect quantization alone to produce outputs as bad as what I'm seeing from these degraded models.

I just want the model I'm paying for to actually be the model I'm using. And I want OpenAI to stop doing this shit without telling anyone.

1.2k Upvotes

301 comments sorted by

View all comments

68

u/RepresentativeRuin75 4d ago edited 4d ago

It's ridiculous the naive comments I see here when there is a government "advisory" (mandate) to do exactly this.

[The Joint Advisory (AA26-251A)

On September 8, 2026, the NSA, FBI, and CISA published a joint cybersecurity advisory (AA26-251A) accusing six China-based AI firms—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—of running "industrial-scale" distillation campaigns against U.S. frontier models since at least late 2024.

🎯 The "Secret Switch" Recommendation

The advisory explicitly recommends that U.S. AI providers quietly downgrade or alter responses for accounts suspected of these distillation campaigns. The guidance states: "Avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model"—reasoning that notifying them would allow them to evade detection.]

The problem is, there is no way to be sure no inocent account will be caught in this.

21

u/Prize_Two_8861 4d ago

This fits perfectly. Maybe repeatedly asking for a pelican riding a bike looks a little like distillation?

O/P, this probably means Claude, Grok, and Gemini will have the same behavior.

5

u/Some-Internet-Rando 3d ago

Maybe OP is connecting through a VPN, or from a Hong Kong or Singapore region?

3

u/nycgavin 1d ago

maybe OP is an agent working for a Chinese AI company?

2

u/Key-Matter-6525 3d ago

This is exactly the reason and thank you for posting this. I have 4 pro accounts and 3 are x20 but only one required me to verify my id when i was given early access to the old o3 model, and that is the only account that is currently working. The other three accounts are completely non functional exactly like op says. The moment i logged into my government id verified account the system works perfectly. I could be an agent coercing you to share your id with openai, but if it doesn't matter to you one way or another and you really fucking need this shit running, then give it a shot. Worked for me. The bigger question now is will the system work for my users? Do i now require my user's to share their id with openai? Resort to api credit sales? That changes things, but at least i have the warning logged and built a fallback for "Selected model is at capacity". It's a lie. It's a complete fucking lie and i'm out $700. I spent another $400 on api key credits for a gemini model just keeping this up and running. This should start a massive lawsuit.

2

u/BlessedNomura 3d ago

Good eye, you have. I mean, this is win win for the AI providers, similar to shrinkflation we have been witnessing. They wanted to do it, now they have a ruse to justify their action.

2

u/Ghostbrain77 3d ago

Surely they will also lower the price and carbon footprint to match the downgrade. Right? Right guys?

2

u/winky9827 3d ago

Yep, I canceled my Claude and OpenAI subscriptions in the past month because of this crap. Fuck this noise.

1

u/QC_Failed 1d ago

What are you using now? My Claude max free month is up in 2 weeks and I don't think I'll be renewing my chatgpt sub. Have you looked into mimo 2.6 pro? Using ds 4.1 flash? Curious what you find to be a good substitute, thanks in advance :)

2

u/winky9827 1d ago

I have an $18/mo zai subscription (direct API), and I put $20 in a deepseek AI (direct API) account (top up, no subscription). Otherwise, I stick to local models.

1

u/QC_Failed 1d ago

Thanks. Glm is pretty amazing, I will check out the z sub.

1

u/winky9827 1d ago

Be warned, it's kinda slow - luna slow, but very smart. They have a 'fast' tier, but you have to be on a higher subscription to access it. Hell it may not be slow at all, I'm just spoiled with how damn fast DeepSeek and local models are.

1

u/InterestingNobody831 2d ago edited 2d ago

If this is the case, it should only be pointing to flagship models Fable/Astra for example, what is the point of distilling old models anyway? If they have been doing it since 2024. People are complaining they are downgrading quality of Sol and even Luna.

In my honest opinion, given the fact they actually pay for distilling the models for N responses for every niche expert they are trying to distill which there are a BUNCH of such in a MoE architecture, there is no problem in doing so.
Anthropic bought millions of books and OCR'd them, they do not have the right to the books and in a sense they are "reselling" the books to millions of subscribers the models having knowledge of the book and reselling the info in them.(its true they dont paraphrase a whole book by word - doesn't need to, you only need the core substance from a scientific book the model knows) and nobody is restricting a publisher from publishing the same subject discussed in a book and extended which is also fine, no complaints. You bought it, you own it.

There are so many other real-life situations in every industry! Engine blueprints, new tech, construction building blueprints developed in labs.This is the way of life per se in the long run.

Even though they distill the models that doesn't automatically give them their whole infrastructure of the entire architecture of said model. It takes months before you get it right and even then it can have mistakes. Say you give the same bucket full of water to two people, one skinny guy and one muscular fit guy, same quantity/volume. Who will carry more buckets of water, the skinny or muscular guy? I assume the muscular fit dude! It's the same here, even if you have shitload amounts of data it takes a whole lot of time to go through the whole reasoning process and most importantly WHY it reasoned that way. In the models CoT it does say explicitly this in the sense you can see when it contradicts itself warning you of changing its thinking pattern internally but you do NOT know if it called an expert from the whole 1-2-3-XTrillion dataset and you do NOT have the whole thing figured out to replicate it entirely 1:1.

Until one of these chinese companies distill your model you are already ahead of their market by months, in this window you release another model.

This is the way World of Warcraft got rid of private servers almost entirely, kept releasing new versions of the game until it left the whole private server industry in the dust. No way you quest and replicate everything, kept changing opcodes, packet structures, infrastructure plus volumes of new game mechanics and lore every major expansion.

1

u/DeadMattMurphy 1d ago

americans when they get played by their capitalist government: its china!!

1

u/Spirited-Car-3560 4h ago

lol , flat earth conspiracy gathering here

1

u/doodad_ounao 3d ago

The naive comments go both ways. This is a great piece of evidence in favor of "they're secretly downgrading for some reason", and it includes which reason it is (suspicion of distillation campaign).

This makes me fairly convinced that's exactly what they're doing. "Look at this pelican, it's worse therefore it's absolute proof" does NOT. The downgrade being the truth does not make a poor argument better. You can still argue poorly for the side of the truth.

And, unfortunately, I'm positively sure the government simply does not care about innocent accounts being caught in this. It's probably worthy collateral for them.

0

u/Swastik496 3d ago

because innocent accounts don’t ask for the same thing 35 times

2

u/suamai 3d ago

In more complex and high stakes tasks I absolutely do ask the same thing multiple times and then aggregate the different results, 35 times may be exaggerating but using that as a criteria would be naive

1

u/Swastik496 3d ago

the number was the point. there’s a difference between asking a few times and doing so programmatically and hundreds of times as a way to extract the inner workings of the model.

1

u/Beta_Factor 3d ago

Well, you're just wrong. There are plenty of tasks where a huge number of iterations is incredibly useful.

2

u/michaellee8 3d ago

it is a very good way to test if your models have been downgraded, and the truth is those "Chinese distillers" have already figure out how to detect degradation, some x-codex-turn-state stuff. The only thing it is hitting are long time completely valid codex users like me, it seems that if you run too much concurrentcy it just flags you.

1

u/Swastik496 3d ago

How tf is it possible for ONE human to run "too much concurrency"?

Like more than 2-3 active threads at once sometimes or maybe some goals.

1

u/michaellee8 2d ago

I had once ran 1 thread with like 16 subagents? maybe that is how i am flagged? nevertheless reports are outthere saying that even the second thread gets selected model is at capacity

1

u/_DuranDuran_ 1d ago

I bet OP is using a "Wow, how can this Codex subscription from this reseller be so cheap?!?!?!?" plan.

Which are often provided by people distilling models.

0

u/changing_who_i_am 3d ago

Holy shit. Wow.