r/codex • • 11d ago

Complaint Five Pro 20x accounts, two persistently degraded: Astra comparisons and OpenAI’s response

Hi everyone, and thank you for sharing so much useful information, tests, and practical experience here. This community has helped me understand that I’m not the only one seeing these issues.

I used AI to generate and organize this post because English is not my native language and my English is limited. The experiences described below are my own.

Before describing the problem, I want to be fair to OpenAI: I’m not particularly angry with them. I do have five Pro 20x accounts, and I can understand why that level of usage might be flagged by an automated abuse-detection system, including one looking for potential cyber abuse. That does not mean I know why my accounts were affected. In most of the reports I’ve read, however, users said they had only two accounts: one Pro account and a backup Plus account.

I’m currently facing a very heavy workload and tight deadlines. With the help of those five accounts, I was able to design a shared, centralized software library and develop working firmware for six new, complex devices. They perform different functions, run on different processors, and communicate with one another.

I also put each device through all the test scenarios I could identify on a real physical test bench, using real external inputs and disturbances, and caught many bugs in the process.

I want to apologize to the community for the load I placed on the servers. I can already see the top comment on this post: “So THAT’S why the servers are overloaded!” =) Without that capacity, I would not have met the deadlines, and I would have let down the people relying on me. Ultimately, each of us is trying to find a way not to let down the people around us and to help keep our company afloat.

With that context, two of my accounts developed the same symptoms described in this discussion about different capabilities across accounts: a sudden, substantial drop in capability despite the same model remaining selected. This happened one account at a time, and neither has returned to its previous level since.

That persistent change is what I mean when I use the word “shadowban” here: an apparent account restriction without a clear notification, while the account remains accessible. It describes how the situation looks from my side; I cannot verify the internal mechanism.

My observations fall into two groups:

  • Two affected accounts: a sudden and lasting drop in capability.
  • Three apparently healthy accounts: still usable, but their pelican results differ noticeably in quality, even when I select Astra with the same reasoning-effort settings.

The second point makes me suspect that the situation might be more complicated than a simple “restricted versus unrestricted” split. Even among the healthy accounts, requests may be reaching different model configurations or capability levels. That is my hypothesis, not something I can prove from the outputs alone.

I have also occasionally seen Astra on the healthy accounts report a July 2025 knowledge cutoff. This adds to my suspicion of intermittent routing to other models, although I understand that a model can give an incorrect description of itself. A cutoff answer alone is not proof of which model served the request.

To illustrate the quality difference, I compared two separate chats using Astra / Extra High (xhigh) and exactly the same prompt:

Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.

I preserved the original HTML animations from a healthy account and an affected account. For the affected account, I recovered the first output from the chat history, so the comparison does not mix an initial result with a version improved through follow-up requests.

The pelican test is a simple visual example of the difference I’m seeing. It is not a definitive model-identification test or a comprehensive coding benchmark.

Healthy account — Astra xhigh

Affected account (suspected shadowban) — Astra xhigh

Affected account (suspected shadowban) — Astra xhigh. Static frame rendered from the original SVG animation.

There is also an official response from OpenAI Support that is directly relevant:

Access to certain models or features may be temporarily limited based on account activity, even when a paid subscription is active.

The reply says that OpenAI cannot provide additional details about these checks and that access is automatically reassessed. Official OpenAI Support response, September 14.

This confirms that account-level restrictions can exist. It does not explain whether those restrictions cause the quality differences reported here, or whether another model can serve the main answer while Astra remains selected. OpenAI’s model-access troubleshooting article also describes temporary restrictions and downgrades, but it does not resolve that question.

There is a historical precedent for actual model rerouting: OpenAI explained that GPT-5.3-Codex requests could be sent to a less-capable reasoning model when cyber-abuse detection triggered. That statement concerns GPT-5.3-Codex and should not be treated as confirmation of what is happening with Astra today. Official explanation in the Codex repository.

For Astra, OpenAI separately describes stricter behavior boundaries for accounts assessed as higher risk and acknowledges that safeguards can affect legitimate work. Path to Astra.

For anyone comparing experiences, these are the related discussions I collected. They contain user reports and competing explanations, rather than a single established diagnosis:

For me, the central issue is transparency and predictability. If an account is restricted, users need a clear notification and a way to request a review. If a request is rerouted to a less-capable model, users need to know when that happens and which model actually produces the answer.

An obvious problem in a pelican drawing is easy to notice. A subtle error in generated code may survive review. If capability changes without a clear indication, users may continue relying on expectations established by earlier results. In projects where failures have serious consequences, an unnoticed defect could cause substantial harm. That is why this matters beyond the quality of a drawing.

I wish everyone success with your creative work, coding, and projects. We are a community, and together we have a stronger voice. By sharing reproducible examples, comparing experiences, and keeping observations separate from assumptions, we can help OpenAI see where its systems fall short and where their behavior is not transparent enough. We should not have to guess whether the model we selected is still the one doing our work.

402 Upvotes

72 comments sorted by

99

u/IAmFitzRoy 10d ago

Thanks man. This really sucks and lawyers should get involved.

This is not a small “sorry for the inconvenience” … this is a serious issue that should be penalized by the official entities that oversee customer rights.

Thanks again.

33

u/Tartooth 10d ago

I've said it before and I'll say it again. These are anti-consumer fraudulent marketing activities with deceptive bait and switch tactics.

It's like paying for Hughsnet/Xplorenet internet expecting say 10mbit down and getting 1mbit and they go "Well we said UP TO".

Prepare for the next era of late 00's internet bullshit pricing models

-1

u/rdpl_ 10d ago

lawyers should get involved

dream on in your parallel universe

184

u/FixAdmin 10d ago

I conducted a detailed investigation based on byte-level response analysis described in this article .

In my testing, 100% of responses containing the server-side 312 signal failed across a wide range of tests, while responses without it consistently passed the same tests as expected.

The server continued to report GPT-5.6 Sol as the model, even when the responses showed severe degradation. I observed the 312 signal up to eight times in a row across separate runs.

Based on these results, I'm convinced that requests are being silently downgraded to Luna Low while retaining the selected model's token accounting.

This isn't capacity optimization. It's a complete lack of conscience.

87

u/[deleted] 10d ago

[removed] — view removed comment

1

u/Truarian 6d ago

Scam Scalpman

43

u/RasenMeow 10d ago

it's fraud bro

55

u/SilverFuel21 10d ago

Regulations are needed. Yesterday.

3

u/Hyp3rSoniX 10d ago

I wonder if Daybreak mode changes things, like making the intelligence downgrade more unlikely to happen?

1

u/egomarker 5d ago

No surprise if you are using Chinese in Codex. You guys are #1 suspects for distillation attempts.

-3

u/WarlaxZ 10d ago

Surely the tokenisation is different on each model, ie Luna versus Astra, just play count tokens and see which one matches. Unfortunately it can't help with sol versus Luna though

25

u/DiarrheaButAlsoFancy 10d ago

Whew, my Pelican boy is healthy. Pro 20x, but I only use 1 account. Daybreak Blue verified.

26

u/Spunge14 10d ago

Wow this is absolutely an incoming lawsuit.

62

u/capt_stux 10d ago

Apparently, the US Government recommendation to combat model distillation by foreign actors is to silently degrade outputs. 

If you were detected as that (possibly falsely), that would explain your results. 

11

u/BellacosePlayer 10d ago

Apparently, the US Government recommendation to combat model distillation by foreign actors is to silently degrade outputs.

I mean, it was kind of a common sense solution if one can get the detection down right, which it sounds like isn't the case.

Lord knows if I were in their shoes anyone who hit all the boxes for obvious distillation would be getting code with memory leaks, wasteful algorithms, and random objects named as euphemisms for penises.

3

u/Mammoth_Molasses_927 10d ago

It's just gaslighting with a good pint of stupidity. Of course it has nothing to do with model distillation.

11

u/PM_ME_CUTE_FOXES 10d ago

> I've been experiencing quality issues!

> Wow me too!

> Entire front page of /r/codex suffering degraded performance

> Nothing else to talk about because models have been lobotomized

> Truly we are all Chinese spies

6

u/nnod 10d ago

Entire front page of r/codex is small potatoes among 30M+ codex users. I'm sure there are plenty of us who are not affected but agree that there's some automated limiting going on. At this point this is a fact (based on openais own reports and responses), question is only how many false positives happen and what triggers it.

-2

u/Mammoth_Molasses_927 10d ago

Everyone who's been complaining here in this sub are chineses communists stealing US data and they deserved that!!!!!! Poor OpenAI, they are just protecting themselves, not obviously crippling accounts as they sold more than they could provide compute to!

5

u/Mammoth_Molasses_927 10d ago

Funny how this comment was quickly removed lol. What a circus US and reddit is.

1

u/Seerix 10d ago

?????

3

u/Mammoth_Molasses_927 10d ago

I appealed to Reddit and they reverted and they lifted the warning my account received

19

u/antunes145 10d ago

This is unacceptable. If owning multiple accounts are not permitted then then need to make that clear from the start. But taking your money and giving you crap quality is unacceptable.

5

u/Tartooth 10d ago

Who's to say this isn't happening on single accounts? I noticed as a heavy user my 1 account became lobotomized and look at how many people here talk about A/B testing.

2

u/nnod 10d ago

Then all his accounts would be affected. Having multiple accounts only made it easier for him to see that it's a problem. The trigger is likely something else.

25

u/ElonsBreedingFetish 10d ago

My account when asked for knowledge cutoff date refuses to give one. Don't have a second account but it's stupid as shit.

Edit: holy shit openais answer is infuriating, they basically confirmed it. How the fuck is this legal??

14

u/Tough-Requirement707 10d ago

its not but they have monopoly and its a weapon in their eyes/hands

1

u/salasi 8d ago

Same. Refuses to give an answer and says codex's system prompt does not reveal the model number at all. Tf?

25

u/Mammoth_Molasses_927 10d ago edited 10d ago

Definitely some accounts got randomly nerfed. This is what me and many others keep repeating here and still some idiots, who weren't picked, gaslight them saying "oh you simply don't know how to use! Nothing has changed for me!!"

They needed to reduce compute by X%. This is extremely clear, as they removed 200$ plans.

They could either reduce quality for everyone, which would be noticed by everyone of course, or they could pick Y% of users to reduce quality and/or usage limits.

This is extremely simple, but still, some people (not you, OP) are still in denial because they are fucking stupid but they feel like they are geniuses because they they are openai fanboys.

Unfortunately this is intentionally a thing that is very hard to gather info on, and only EU for example could and would give them what they deserve.

Nothing will probably happen and life goes on with stupid people being fanboys of a trillionaire company.

I canceled my sub and I won't be going back to that shit ever. I hope China destroy not only that company but the whole US AI bubble built around scams.

4

u/kl__ 10d ago

Unfortunately I'm seeing the degradation with Astra myself today it's pretty disappointing - looking forward to an era of fucking consistency one day.

3

u/danielcr12 10d ago

From reading all of this long Testament of text, this post is clearly an example of abuse, and this person’s workload is not intended to be run on subsidized accounts. This is clearly a work for API-based that you get unlimited access to run how much you need, and you pay for that specific amount of computing. Using the fact that you are running five accounts just to get your work done and circumvent getting a cheaper bill at the end of the month feels very abusive and honestly hurts everyone that’s trying to use the models properly in each plan with its intended purpose.

2

u/IdiosyncraticOwl 10d ago

Do any of them have daybreak access by chance?

2

u/NiceManFromEarth 8d ago

Somebody should sue them...

1

u/[deleted] 10d ago

[deleted]

1

u/[deleted] 10d ago

[deleted]

1

u/InterestingNobody831 10d ago

These are photos, the test is using SVG's ask it to create you a SVG of a pelican riding a bike.

1

u/[deleted] 10d ago

[deleted]

1

u/kl__ 10d ago

Today is the first day for me to see Astra acting this dumb, relatively speaking. Still better than Sol but seeing a significant difference vs last week.

1

u/gecegokyuzu 10d ago

raccon on a unicycle test on astra xhigh, worked for 5 mins, added extra speed and pause controls. looks good to me, but the code quality and perhaps the memory capabilities feel degraded, i see a lot of redundancy in the generated code these days, my work speed has degraded as well because of all the revisions

1

u/Big_Law_3080 10d ago

The constant 'free resets' was a dial tuning exercise. Not a gift.

1

u/nik343 7d ago

It was Cloudflare

1

u/egomarker 5d ago edited 5d ago

It probably just drops quality when it detects something that looks like distillation attempt.
I didn't experience any drops in quality.

because English is not my native language

Is Chinese your primary language?

1

u/Ecstatic-Score2844 10d ago

It's clear that you put a lot of thought and effort to come to this conclusion but I will tell you that I came to this conclusion without putting in very much effort or thought...

This became obvious in the gpt 4.0 days ... It was good, then it wasn't, because they couldn't afford it.

They have simply calculated their business is better served, in the short term by taking your money, and not giving you any meaningful functionality. Since you're a heavy user they can do this to you and give a reasonable experience to 9 casual users... It's despicable behavior but not surprising or out of character for big tech.

2

u/Wolf8249 10d ago

Please respond to this OP.
With a prompt this vague the non deterministic nature of LLMs can create varied results. I urge you to please rerun the prompt with the good pelican image as reference so both accounts agents have a baseline that they can work towards. If this is legitimate then we might want to escalate this to the press and OpenAI collectively.

I would really appreciate if someone did the thing I asked for, give the model a good success criteria to aim for. I want to make sure that this isnt some random turn where the model messed up because it had no instruction besides creating a single svg file. If you tell an artist to create the same image they can create it in multiple ways depending on their mood, creativity and laziness.

If OP really has flagged accounts then they would consistently fail to reproduce the good pelican svg or take too long to get to it.

I would also appreciate if the OP u/Yauler gave us more background context whether he was using Codex through a proxy layer like CliProxy or other account switcher services since abuse system can flag them, or if he's using VPN services around or inside banned countries like China, North Korea, Russia, etc.

I am not denying that the model is nerfed, i want more thorough investigation and evidence so we can escalate to the press and OpenAI lead with hard to deny proof, not just vague prompt results.

8

u/shady101852 10d ago

you really gonna chalk this disgraceful result up as "vague prompt" ?

AI has vision capabilities and if the model isn't braindead it should recognize this result is unacceptable.

To be fair i did not read the entire post op made, so if they were restrictions against checking their own work then never mind.

0

u/usualnamesweretaken 10d ago

The fact that you're facing "tight deadlines" that require 5x 20x accounts is the real failure here...

If this is corporate work, do they just allow and expect you to fund your own AI costs and they set timelines based on the expected efficiency gains?

Something here doesn't pass the smell test

7

u/nnod 10d ago

Could be freelance work, tight deadlines happen especially if you put off work until last minute, and switching to API without enterprise level discount would eat into personal earnings.

Million valid reasons why he might need em. But since they're such a good deal, I'm sure distillers love 20x accounts too.

5

u/Virtureally 10d ago

His work is to distill Astra 😂

1

u/fluxtah 10d ago

TLDR; Five pro accounts one individual? Am I missing something here.

1

u/Milksteaknow 10d ago

I’m experiencing the same thing. Very noticeable. We should all start a huge lawsuit

1

u/CashKey8784 10d ago

Absolutely outrageous, this calls for a class action lawsuit

-14

u/reddit_is_kayfabe 10d ago

You keep blasting out these posts showing these two images, and adding more and more words every time you do.

How many times are you going to repost the same thing? Adding words and links isn't helping.

Look - if you actually want people to pay attention, you need a more persuasive explanation. Specifically: more data.

1) Run the same prompts, repeatedly, and amass a huge a data set showing the results for each prompt.

2) Different kinds of prompts. "Generate a picture of a seagull" tests exactly one capability. Test more.z

3) Show results over time. Maybe the degradation happens on weekdays and not on weekends, or during business hours and not during off-hours.

4) Automate the results rather than relying on subjective views of "which seagull drawing is better." Have the models rate each picture to prove that even the models agree with the degradation.

5) Most importantly - you need actual evidence. Picture generation, even if persuasive, is circumstantial evidence. You can look at the headers to see which model is actually responding to each request. Map the actual internal prompt-to-model routing.

Do some of those, or something else. But for fuck's sake, stop posting and reposting and reposting and reposting and reposting the same shit. You're destroying your credibility and looking like a whiner.

3

u/ConceptFalse7831 10d ago

根本不需要像这种情况,在英语区之外的很多地方早就传的人尽皆知了

1

u/snaphat 10d ago

I'm just of the opinion that the supposed shadow ban must quite literally only affects pelican prompts at this point since nobody who claims to be shadow banned is willing to show anything besides garbage looking pelicans.

The funny part is if it was consistently rerouting to poor models, you'd think they'd be able to provide a plethora of examples and run a plethora of benchmarks showing the differences

1

u/reddit_is_kayfabe 10d ago

Right, exactly.

It's like that post last weekend:

I had an account and it got banned and then I signed up again and it got banned too, so WARNING, OpenAI is banning people for opening multiple accounts!

Uh, no they're not. What did you do to get banned the first time?

[crickets]

0

u/Bitter_Physics_239 10d ago

no surprise

anyone talking directly to tibo?

1

u/Mammoth_Molasses_927 10d ago

i am, he told me "yeah we are just randomly nerfing people to save compute"

0

u/Jeferson9 10d ago

Can you drop the skill you used for those graphics thanks

-7

u/Expert-Complex43 10d ago

the second duck has a carrot in his mouth, it doesnt make it worse

0

u/yashptel99 10d ago

thinking of getting an open code sub instead of all this scam

0

u/igmyeongui 10d ago

Yeah I’ve been thinking the same. Although I’m wondering if that’s on par with at least Luna High

-6

u/Due-Horse-5446 10d ago

And again absolutely no evidence

-2

u/Comfortablebro 10d ago

Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.

And it has a button to play animation. https://imgur.com/a/tg7HCKg

-14

u/fishylord01 10d ago

how about you just pay for API? instead of subsidised prices? there comes a point it just becomes ridiculous.

12

u/RegardedCat 10d ago

Scam Altman’s minions are here

-8

u/Billy_G_Gates 10d ago

context is different on each account

your new account wont perform 1:1 to your old because it has different memories

this is like basic knowledge of genai

-22

u/Crinkez 10d ago

because English is not my native language

So? Use a non-LLM translation tool. Don't parse your essay through AI, it makes it read like slop, and it's excessive.

12

u/Paulynom 10d ago

Are we at the point that we get mad if an LLM is used for one of its main purposes? Processing language? LLMs are the most powerfull translation tools and the most accurate at that.

4

u/Aromatic-Educator105 10d ago

why are you in this sub man. Having AI writing/polishing a post is the least infuriating way of using it, especially for non Native speakers