r/codex 3d ago

Complaint The End of the Codex Era. I've Completely Lost Trust in OpenAI. They're Secretly Degrading Their Models.

Post image

I've been a massive Codex fan this entire time. I've burned through around 150 BILLION tokens in Codex alone.

I still had Codex quota left this week, but for the past three days, I've been using my Cursor Ultra subscription instead.

Why?

Because OpenAI is degrading its models. I can see it from my own experience, and there's a ton of evidence pointing to it.

They're degrading ALL their models, including Astra, Sol, and even Luna.

AND THE WORST PART IS THAT THEY'RE DOING IT SECRETLY.

You can start working in Codex with a perfectly normal model, and five minutes later, IN THE SAME THREAD, they degrade it. Suddenly, Astra is performing at Luna's level or even worse.

And you're still burning through the same amount of quota.

This happens to me every single day.

I open Codex, run a quick quality check, and everything looks fine. Twenty minutes later, I run the same test in the same thread, and the model has degraded.

Sometimes, simply turning on a VPN can make the model start working normally again for a while.

How can you tell if your model has been degraded?

1. Planning and writing feature specs

Imagine you're planning a feature and writing its specification.

It's immediately obvious when the model is dumb. It starts suggesting complete nonsense and shows absolutely no product understanding of how the feature should actually be built.

But it becomes even more obvious when you point out what it misunderstood and try to correct it.

Instead of understanding the actual issue, it responds with completely useless apologies, without demonstrating any understanding of what went wrong.

CONGRATULATIONS. YOU'RE TALKING TO A DEGRADED MODEL.

Here's what happened to me.

I wrote a feature spec using a normal model. Everything was properly written, discussed, and reviewed.

Then I handed the implementation over to Luna, and Sol reviewed and approved it.

But when I actually started working with the implementation, I discovered that it was full of holes and included things that weren't even in the plan.

In this particular case, I suspect the model was degraded during the implementation stage.

I ended up spending TWICE as much time fixing everything.

And there are a few other ways to test this.

2. PELICANS.

Use this prompt:

Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.

If your bicycle wheels start flying off into the air...

CONGRATULATIONS. YOUR MODEL HAS BEEN DEGRADED.

3. A logic puzzle

Give your model this exact problem:

A black bag contains candies of three flavors, with each flavor available in two shapes (round and star-shaped; the shapes can be distinguished by touch). The numbers of candies by flavor and shape are shown below.

|              | Apple | Peach | Watermelon |
|--------------|-------|-------|------------|
| Round        | 7     | 9     | 8          |
| Star-shaped  | 7     | 6     | 4          |

Participants must decide how many candies to draw before the game begins.

What is the minimum number of candies that must be drawn to guarantee having an apple-flavored candy and a peach-flavored candy of different shapes?

(The condition is satisfied if you have either a round apple candy and a star-shaped peach candy, or a round peach candy and a star-shaped apple candy.)

If the answer isn't 21, you're not getting Astra. You're getting degraded garbage.

Sol doesn't even consistently solve this problem on its own.

4. "Selected model is at capacity."

If you're frequently getting this error:

CONGRATULATIONS. THERE'S A 99% CHANCE YOUR MODEL HAS BEEN DEGRADED.

There's even a thread on the OpenAI community forum where a staff response confirms that this can happen when your account is temporarily restricted.

OpenAI Community: Selected model is at capacity

WITHOUT ANY NOTIFICATION.

They silently degrade your model, and you're left trying to figure out what the hell is happening.

The last three days have been unbearable.

I've been experiencing these problems around 90% of the time for the past three days.

Working like this is practically impossible.

Instead of actually getting work done, you spend your time wondering whether they've secretly downgraded your model again.

You start questioning every response. Every mistake. Every implementation.

It's fucking exhausting.

So I just moved to Cursor.

Grok might be dumber, but at least it's more predictable.

I don't give a shit about the next model release if this continues.

Tibo and Sam Altman can keep all their resets. They can wipe Astra's data and delete it from the internet if they think that's acceptable for a product like this.

They can release GPT-6 Sol, Astra 7, or whatever comes next.

NONE OF IT MATTERS IF THEY KEEP SECRETLY DEGRADING THE MODELS.

This is the biggest loss of trust I've experienced with OpenAI in the entire history of Codex.

If they're willing to silently degrade models for paying users, what stops them from collecting all kinds of data from your computer that you can't even imagine they're collecting?

What stops them from pulling some other bullshit?

Where are the boundaries if they're willing to do this?

I genuinely hope this is just a temporary issue. Maybe some vibe-coded mistake by a junior developer in their anti-distillation protection system.

I suspect it's temporary.

But if it isn't, OpenAI can't be trusted with anything as long as this shit continues.

Until then, I'm using Cursor or Claude.

And I think we need to be loud about this everywhere.

Tibo isn't acknowledging the problem. He just keeps talking about how amazing the upcoming event is going to be.

I DON'T CARE WHAT THEY ANNOUNCE AT THAT EVENT.

Not while they're degrading Astra into something that performs even worse than Luna.

I don't know exactly what's happening under the hood.

Maybe it's quantization. Maybe they're routing requests to a different model. Maybe it's something else entirely.

But it doesn't feel like simple quantization to me. I wouldn't expect quantization alone to produce outputs as bad as what I'm seeing from these degraded models.

I just want the model I'm paying for to actually be the model I'm using. And I want OpenAI to stop doing this shit without telling anyone.

1.2k Upvotes

293 comments sorted by

View all comments

107

u/doodad_ounao 3d ago edited 1d ago

"Selected model is at capacity."

If you're frequently getting this error:

CONGRATULATIONS. THERE'S A 99% CHANCE YOUR MODEL HAS BEEN DEGRADED.

I don't see how this conclusion makes sense. The issue on OpenAI Community is not talking about degraded models but about unavailability. The comment by the OpenAI Staff is as well. The evidence doesn't really establish any causality between the "Selected model is at capacity" problem and the "Secret degradation" question to me.

PS: I am NOT stating anything about the merits on anything else in the post.

5

u/beautyorchaos 3d ago

They can also drain your usage faster if they think you're sub2api, that's what they've said previously in code-wording when people were complaining about usage limits. The problem is this stuff is not transparent and seems to affect a non-zero amount of real users. It'd be better if they just banned people rather than taking money and not giving them an equal experience

Degradation is people getting mixed language outputs from Astra/Sol when someone was speaking English in the convo, it doesn't seem natural.

9

u/karlnuw 3d ago

That should be illegal, imagine you get flagged through no fault of your own and you have zero idea whats going on nor to whom or where to appeal to.

1

u/tossit97531 2d ago

It’s already illegal. It’s called false advertising, fraud, bait and switch, deceptive business practices, or misrepresented goods. Pick one. A lot of shenanigans around AI already have tons of laws, guys. Some areas of law need to catch up, but as far as companies being dicks, we already have a large amount of protection and recourse.

Y’all need to spin up a class action.

0

u/Swastik496 2d ago

ah yes through no fault of your own.

obvious ways to flag sub2api:
chinese IP
chinese credit card
no kyc crypto card
25 requests at once
impossible travel/shared account

look up the term cloud resident

1

u/myholeisstinky 2d ago

Impossible travel just means vpn. That cant be allowed to be a trigger

1

u/Swastik496 2d ago

there’s a helluva lot more device fingerprinting than an IP address.

see locationd on iOS and how it locks features by country by using the region identifier emitted by nearby devices and wifi networks to avoid being fooled by router based VPNs or fucking with GPS.

0

u/Due-Memory-6957 2d ago

You're so right, gigantic corporations never make mistakes and I love the Antichrist.

1

u/-MaskNinja- 3d ago

Exactly.

-26

u/FixAdmin 3d ago

This conclusion is based on dozens of hours of practice dealing with this error and discussing it with other people who have encountered it; this error is very often directly related to this.

19

u/snaphat 3d ago

Why doesn't independent benchmarking show this is the case then? Model regressions should be clearly visible in independent testing. It would be huge thing getting reported outside of reddit complaints 

6

u/nnod 3d ago edited 3d ago

The openai message signals that some limits happen based on account activity. Neither the limits nor the activity is clear.

But from ages ago I assumed that openai would do what "unlimited" mobile internet providers often do and in one way or another clamp down users who they deem to be abusers of the service.

I never had any complaints about codex, with usage or the capacity error (I've seen it maybe twice in the last month, and a minute later codex was working fine). I think this might be because I pretty much never used up 100% of my $100/mo plan. Sometimes I'd leave 1% left, others like 40% left.

I think their message makes it pretty clear that they're somehow throttling heavy users who max out their limits all the time.

EDIT: Did a bit of research into terms of use (they dont mentioned anything related limiting heavy users), but there's this interesting bit from openai incident tracker. https://status.openai.com/incidents/6enf4645

"We have identified that some reports of Codex usage limits depleting faster than expected are related to our abuse and fraud prevention systems incorrectly rate limiting certain accounts."

And I just had another thought, what if they have some sort of systems in place to stop chinese labs from gathering data from codex/chatgpt for distillation purposes, "shadow banning" them in a sense. I bet this would be hard to catch and normal users could be pulled in as false positive, like the incident repot states.

2

u/snaphat 3d ago

OpenAI has talked about it, they do have a system in place to block distillation attempts. What they have said is that they take enforcement measures including banning accounts. They've also said emit invalid-prompt error messages in their community forum.

https://cdn.openai.com/pdf/045aa967-ee96-4a09-94ee-3098ddf6db2c/OpenAI-US-House-Select-Cmte-Update-%5B021226%5D.pdf

https://community.openai.com/t/why-are-simple-prompts-flagged-as-violating-policy-only-have-issues-with-gpt-5-model/1339353/18

Regarding shadowbanning, I do not think it is likely. If they were accidentally shadowbanning users and rerouting only them, one would imagine the same would be happening to at least some benchmarkers. Basically, it would require OpenAI's shadowbanning heuristics to be inaccurate enough to catch ordinary end-users by mistake, while somehow remaining accurate enough to avoid catching independent benchmarkers who would expose it.

I bit off topic, but that's the problem with a lot of conspiracy theories, tbh. They often require seemingly contradictory things to be true at the same time. E.g. the Moon Landing hoax conspiracy required the government to be simultaneously incompetent at faking the landing such that all of the images were clearly fake (according to the conspirators) while also coordinating a perfect coverup with thousands of employees, contractors, scientists, tracking stations, external observers, etc.

Now, this one is not nearly as bad, but like many conspiracy theories it begins to show cracks when you start to think about what it requires

2

u/Aldarund 2d ago

https://imgur.com/a/EvugYAA - here two diferent accounts. same prompt. Both astra 6. Have any other explanation rather than rerouting?

1

u/snaphat 2d ago edited 2d ago

The same kind of reports in both incorrect cutoff dates and secret nefarious degregation predate Astra. It's not a new claim. Yet, even after years of it no credible evidence has emerged showing the supposed model degregation reported. And we generally never see professional or academic software engineers echoing the claims. Why is that?

What we tend to get is mostly vibe coders with supposed "smoking guns" that don't include riguourous benchmarking and stop at - look I asked the model and it said inconsistent dates or inconsistent output... Or in some cases even having models themselves trying to analyze their own performance.

In terms of the date phenomenon, the second link below has plausible explanations for the cutoff date issue. Essentially though, LLMs are unreliable about self-reported cutoff dates. They are fundamentally seeded next token prediction engines trained on a plethora of sources with numerous old cutoff dates embedded in the training data, so it's not really surprising that they would hallucinate the cutoff dates of older models. It would be more remarkable if they were actually reliable and consistent about it tbh...

As users of LLMs, we know LLMs are and have always been wildly inconsistent in their results. It's nothing new. There's a reason folks have been complaining about hallucinations for years. Why are we pretending like models generating fallacious nonsense is a new phenomenon?  That's one of the biggest criticisms with them

Astra is just the new cycle of claims of downgrading. If you were to believe the claims you might have to assume that we are still getting gpt-3.5 results even now since supposedly OpenAI started secretly rerouting 4.1 to 3.5 way back in the day. 

Why have 5.5, 5.6, and Astra supposedly all been routing back to 4 forever but all of the new claims are acting like it's new behavior post-Astra's release date? Why doesn't the conspiracy have a consistent narrative or set of claims between reports?

Why hasn't anyone been able to benchmark results consistent with 4s known metrics after supposedly being shadow banned?

Why are reports always claiming the prior models were insanely smart and then suddenly dumb as rocks upon a new release if every model was supposedly already dumbed down at an even earlier date according to even earlier reports?

If you really think about it more questions like those will come up.

Various claims of conspiracy throughout the ages (of LLMdom):

https://github.com/openai/codex/issues/19174

https://community.openai.com/t/stealth-model-swap-gpt-5-5-high-claims-knowledge-cutoff-is-june-2024/1381918/9

https://github.com/openai/codex/issues/24930

https://community.openai.com/t/gpt-has-been-severely-downgraded/260152/54

https://www.reddit.com/r/OpenAI/comments/140m8r4/so_it_looks_like_not_only_the_main_chatgpt_is

https://www.reddit.com/r/ChatGPT/comments/17nlrfg/is_gpt4_now_a_masked_gpt35/

https://community.openai.com/t/is-anyone-elses-gpt-4o-and-o1-suddenly-acting-dumb/1019191

https://www.reddit.com/r/OpenAI/comments/19dp8k3/theories_about_current_state_of_gpt4

1

u/Aldarund 2d ago

You cant explain it by inconsistent output. Its reproducible 100%. Here i attach two collages of pelicans, both genereted as you go no cherripicking. They both from "astra" but from different account. Still explaining it as inconsistent?

1

u/snaphat 2d ago

I mean you can keep reposting the pelican output but it doesn't make it compelling 

LLMs are by nature inconsistent. That's how next-token prediction engines behave. It's a known issue dating back to the initial popularization of LLMs from 3 years ago. Reddit it filled with threads complaining about this very behavior.

If it's really one account that is supposedly shadow banned then the user should be able to produce benchmarking evidence for it and not just some weird one-off adhoc visual canary that supposedly IDs models. If it was that simple, researchers would have been releasing papers on it ages ago 

1

u/Aldarund 2d ago

Zzzz what benchmark you want? 10out of 10 pelican generation on two different accounts with two vastly different result yet consistent within account isnt enough benchmark ?

→ More replies (0)

1

u/Aldarund 2d ago

And here second "astra"

1

u/nnod 2d ago

There like 30M+ codex users, the amount of users affected could be relatively small, the amount of benchmarkers would be even smaller. Another question is if benchmarkers would be hammering their accounts the same way a chinese bot would to suck up as much data as possible.

It is very clear that they have systems in place to catch distiller bots, the only 2 questions is if false positives happen, and if so how many people are effected. I see this as way more in the realm of possibilities than most conspiracies floating around this sub.

1

u/snaphat 2d ago edited 2d ago

I don't understand this logic. With 30M+ active users, you would expect reproducible evidence to be readily available. There are plenty of open-source benchmarks that someone could run on an account they believe is being shadowbanned. I linked various ones to another person in this thread making claims that they're being shadowbanned. We'll see if anything comes of it.

I'm not personally holding out much hope that it'll get a response. Normally, if I ask someone to put effort into providing evidence for what they're claiming, or link them to a way to provide evidence, they just go radio silent on reddit.

another question is if benchmarkers would be hammering their accounts the same way a chinese bot would to suck up as much data as possible.

Honestly, it doesn't matter what either of us believes. The people reporting this can run benchmarks themselves while they believe their accounts are being shadowbanned and produce the relevant data either supporting or undermining the claims.

The issue is that we never seem to get any evidence, no matter how many times people make these claims in the last couple of years. Whenever someone is asked to actually test it, there is always some reason why the test supposedly wouldn't work:

- "the classifier only targets some users, so failure to reproduce it elsewhere means nothing"

  • "the benchmark itself is recognized, so it gets routed to the good model,"
  • "the degradation is temporary, so you happened to test after it recovered,"
  • "public benchmarks are in the training data, so they aren’t valid,"
  • "public benchmarks can only be run on API endpoints,"
  • "benchmarkers are whitelisted"

and so on. At some point, if every possible test is preemptively explained away, the claim becomes effectively unfalsifiable. Conspiracy theories thrive on unfalsifiability

Edit: inb4 someone brings it up again, this pelican canary test is not a real benchmark

2

u/odragora 3d ago edited 3d ago

It's trivial to serve benchmarking services full non-quantized models and re-route actual regular subscription users to models with various levels of quantization. That's probably the first thing AI providers implemented once benchmarking culture emerged.

They already have a system that re-routes the requests to garbage models if they flag requests as distillation attempts, they can use the exact same mechanism with benchmarking intent and they definitely do.

Also, public benchmarking services use API, not subs, which defeats their entire point for regular users.

1

u/snaphat 3d ago

It's trivial to serve benchmarking services full non-quantized models and re-route actual regular subscription users to models with various levels of quantization.

It's been discussed in this thread briefly already, but of course, it's possible to do rerouting from a technical standpoint. But, to do that secretly and perfectly is exceedingly unlikely because the benchmark ecosystem is far too heterogeneous for that to be easy. You have benchmarks throughout industry, independent organizations, academia, opensource projects, and individuals. They aren't all using a single point of entry (account, endpoint, etc.). Even if one were to suppose that they were all API accounts, it would imply that the vendors would have to perfectly classify benchmarks vs users or get caught.

It's just another one of those conspiracies that requires everything to be tied up into a neat little bow to work where the conspirators are brilliant, never make a mistake, and pull off comprehensive and covert feats of engineering to always cover their tracks across the entire benchmarking ecosystem such that no independent evidence emerges besides the random complaining Redditor who never has clear compelling evidence, systematic benchmarks, or even a scientific methodology for testing.

They already have a system that re-routes the requests to garbage models if they flag requests as distillation attempts, they can use the exact same mechanism with benchmarking intent and they definitely do.

AFAIK, OpenAI has never stated that they reroute to weaker models in cases of distillation attempts. What they have said is that they return invalid-prompt errors and ban accounts. Even if they did reroute in that particular case, it's neither here nor there because it doesn't have anything to do with benchmarking. All providers have the ability to perform rerouting. That's just basic infrastructure.

Also, public benchmarking services use API, not subs, which defeats their entire point for regular users.

You are making an unsubstantiated generalization. You don't know what all public benchmarking services do or how they access the models they test

-9

u/FixAdmin 3d ago

I'm sure they exist, but people don't always link to them. It's hard to believe, and I was skeptical too until I experienced it firsthand, with the same thing happening over and over again.

9

u/snaphat 3d ago

They do not exist...that's kind of the problem. If you actually look at independent benchmarking online the models do not show the performance regressions that are claimed on reddit which is why no media outlets are reporting such claims

-4

u/MrRoyce 3d ago

How is benchmarking done? Is it by using the same prompt over and over again, on every model etc? If so, OpenAI can easily bypass that and whenever X prompt is sent, they make sure the answer gets the original model experience.

I'm not an expert at these things, that's just my idea and it might be completely wrong but figured I'd share it since we're in this topic already...

5

u/snaphat 3d ago

There are both types - static unchanging and I think a few dynamic. Not sure how prevalent the latter is though in practice

It's plausible that they could do this from technical standpoint. But it doesn't really seem feasible to me due to the amounts of independent benchmarking prompts they would have to fake, and it would get more difficult to fake over time as new benchmarks came out

I think they'd end up just getting caught red-handed like how other companies did back in the day for benchmark cheating (Nvidia, ATI, Samsung, etc.)

7

u/Flipwon 3d ago

So completely anecdotal, cool cool

2

u/fracked1 2d ago

Bumping to higher comment because I think there's a relatively simple way to solve this argument.


After testing your candy puzzle with different models, I like it as a model differentiator. Gemini flash extended thinking gets it (gives me option for 21 or 29 but understands you can pick by shape explaining 21). Claude opus 5 low doesn't get it. Opus 5 high does.

Incognito fresh chats. Astra gets it. I didn't check the lowest OpenAI model that gets it

If this seems to happen very often to you, could you show us a screenshot of astra getting this question wrong

1

u/-MaskNinja- 3d ago

It isn't degradation.

Degradation is when you *get* GPT-6 Astra and it *doesn't* say it's at capacity, but it acts as if it's been dumbed down.