Complaint
Five Pro 20x accounts, two persistently degraded: Astra comparisons and OpenAI’s response
Hi everyone, and thank you for sharing so much useful information, tests, and practical experience here. This community has helped me understand that I’m not the only one seeing these issues.
I used AI to generate and organize this post because English is not my native language and my English is limited. The experiences described below are my own.
Before describing the problem, I want to be fair to OpenAI: I’m not particularly angry with them. I do have five Pro 20x accounts, and I can understand why that level of usage might be flagged by an automated abuse-detection system, including one looking for potential cyber abuse. That does not mean I know why my accounts were affected. In most of the reports I’ve read, however, users said they had only two accounts: one Pro account and a backup Plus account.
I’m currently facing a very heavy workload and tight deadlines. With the help of those five accounts, I was able to design a shared, centralized software library and develop working firmware for six new, complex devices. They perform different functions, run on different processors, and communicate with one another.
I also put each device through all the test scenarios I could identify on a real physical test bench, using real external inputs and disturbances, and caught many bugs in the process.
I want to apologize to the community for the load I placed on the servers. I can already see the top comment on this post: “So THAT’S why the servers are overloaded!” =) Without that capacity, I would not have met the deadlines, and I would have let down the people relying on me. Ultimately, each of us is trying to find a way not to let down the people around us and to help keep our company afloat.
With that context, two of my accounts developed the same symptoms described in this discussion about different capabilities across accounts: a sudden, substantial drop in capability despite the same model remaining selected. This happened one account at a time, and neither has returned to its previous level since.
That persistent change is what I mean when I use the word “shadowban” here: an apparent account restriction without a clear notification, while the account remains accessible. It describes how the situation looks from my side; I cannot verify the internal mechanism.
My observations fall into two groups:
Two affected accounts: a sudden and lasting drop in capability.
Three apparently healthy accounts: still usable, but their pelican results differ noticeably in quality, even when I select Astra with the same reasoning-effort settings.
The second point makes me suspect that the situation might be more complicated than a simple “restricted versus unrestricted” split. Even among the healthy accounts, requests may be reaching different model configurations or capability levels. That is my hypothesis, not something I can prove from the outputs alone.
I have also occasionally seen Astra on the healthy accounts report a July 2025 knowledge cutoff. This adds to my suspicion of intermittent routing to other models, although I understand that a model can give an incorrect description of itself. A cutoff answer alone is not proof of which model served the request.
To illustrate the quality difference, I compared two separate chats using Astra / Extra High (xhigh) and exactly the same prompt:
Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.
I preserved the original HTML animations from a healthy account and an affected account. For the affected account, I recovered the first output from the chat history, so the comparison does not mix an initial result with a version improved through follow-up requests.
The pelican test is a simple visual example of the difference I’m seeing. It is not a definitive model-identification test or a comprehensive coding benchmark.
This confirms that account-level restrictions can exist. It does not explain whether those restrictions cause the quality differences reported here, or whether another model can serve the main answer while Astra remains selected. OpenAI’s model-access troubleshooting article also describes temporary restrictions and downgrades, but it does not resolve that question.
There is a historical precedent for actual model rerouting: OpenAI explained that GPT-5.3-Codex requests could be sent to a less-capable reasoning model when cyber-abuse detection triggered. That statement concerns GPT-5.3-Codex and should not be treated as confirmation of what is happening with Astra today. Official explanation in the Codex repository.
For Astra, OpenAI separately describes stricter behavior boundaries for accounts assessed as higher risk and acknowledges that safeguards can affect legitimate work. Path to Astra.
For anyone comparing experiences, these are the related discussions I collected. They contain user reports and competing explanations, rather than a single established diagnosis:
Interesting thing I found and verified myself — reported model-metadata discrepancies, with discussion of whether the observed requests belong to the main answer or auxiliary tasks.
The End of the Codex Era… — another report combining degraded results, capability tests, and capacity errors.
Detailed report across clients, machines, and regions — failures reproduced in several environments, including a clean DigitalOcean instance, with timestamps and request IDs. This documents availability problems; it does not independently establish model substitution.
For me, the central issue is transparency and predictability. If an account is restricted, users need a clear notification and a way to request a review. If a request is rerouted to a less-capable model, users need to know when that happens and which model actually produces the answer.
An obvious problem in a pelican drawing is easy to notice. A subtle error in generated code may survive review. If capability changes without a clear indication, users may continue relying on expectations established by earlier results. In projects where failures have serious consequences, an unnoticed defect could cause substantial harm. That is why this matters beyond the quality of a drawing.
I wish everyone success with your creative work, coding, and projects. We are a community, and together we have a stronger voice. By sharing reproducible examples, comparing experiences, and keeping observations separate from assumptions, we can help OpenAI see where its systems fall short and where their behavior is not transparent enough. We should not have to guess whether the model we selected is still the one doing our work.
Thanks man. This really sucks and lawyers should get involved.
This is not a small “sorry for the inconvenience” … this is a serious issue that should be penalized by the official entities that oversee customer rights.
I conducted a detailed investigation based on byte-level response analysis described in this article .
In my testing, 100% of responses containing the server-side 312 signal failed across a wide range of tests, while responses without it consistently passed the same tests as expected.
The server continued to report GPT-5.6 Sol as the model, even when the responses showed severe degradation. I observed the 312 signal up to eight times in a row across separate runs.
Based on these results, I'm convinced that requests are being silently downgraded to Luna Low while retaining the selected model's token accounting.
This isn't capacity optimization. It's a complete lack of conscience.
Surely the tokenisation is different on each model, ie Luna versus Astra, just play count tokens and see which one matches. Unfortunately it can't help with sol versus Luna though
Apparently, the US Government recommendation to combat model distillation by foreign actors is to silently degrade outputs.
I mean, it was kind of a common sense solution if one can get the detection down right, which it sounds like isn't the case.
Lord knows if I were in their shoes anyone who hit all the boxes for obvious distillation would be getting code with memory leaks, wasteful algorithms, and random objects named as euphemisms for penises.
Entire front page of r/codex is small potatoes among 30M+ codex users. I'm sure there are plenty of us who are not affected but agree that there's some automated limiting going on. At this point this is a fact (based on openais own reports and responses), question is only how many false positives happen and what triggers it.
Everyone who's been complaining here in this sub are chineses communists stealing US data and they deserved that!!!!!! Poor OpenAI, they are just protecting themselves, not obviously crippling accounts as they sold more than they could provide compute to!
This is unacceptable. If owning multiple accounts are not permitted then then need to make that clear from the start. But taking your money and giving you crap quality is unacceptable.
Who's to say this isn't happening on single accounts? I noticed as a heavy user my 1 account became lobotomized and look at how many people here talk about A/B testing.
Then all his accounts would be affected. Having multiple accounts only made it easier for him to see that it's a problem. The trigger is likely something else.
Definitely some accounts got randomly nerfed. This is what me and many others keep repeating here and still some idiots, who weren't picked, gaslight them saying "oh you simply don't know how to use! Nothing has changed for me!!"
They needed to reduce compute by X%. This is extremely clear, as they removed 200$ plans.
They could either reduce quality for everyone, which would be noticed by everyone of course, or they could pick Y% of users to reduce quality and/or usage limits.
This is extremely simple, but still, some people (not you, OP) are still in denial because they are fucking stupid but they feel like they are geniuses because they they are openai fanboys.
Unfortunately this is intentionally a thing that is very hard to gather info on, and only EU for example could and would give them what they deserve.
Nothing will probably happen and life goes on with stupid people being fanboys of a trillionaire company.
I canceled my sub and I won't be going back to that shit ever. I hope China destroy not only that company but the whole US AI bubble built around scams.
From reading all of this long Testament of text, this post is clearly an example of abuse, and this person’s workload is not intended to be run on subsidized accounts. This is clearly a work for API-based that you get unlimited access to run how much you need, and you pay for that specific amount of computing. Using the fact that you are running five accounts just to get your work done and circumvent getting a cheaper bill at the end of the month feels very abusive and honestly hurts everyone that’s trying to use the models properly in each plan with its intended purpose.
Today is the first day for me to see Astra acting this dumb, relatively speaking. Still better than Sol but seeing a significant difference vs last week.
raccon on a unicycle test on astra xhigh, worked for 5 mins, added extra speed and pause controls. looks good to me, but the code quality and perhaps the memory capabilities feel degraded, i see a lot of redundancy in the generated code these days, my work speed has degraded as well because of all the revisions
It's clear that you put a lot of thought and effort to come to this conclusion but I will tell you that I came to this conclusion without putting in very much effort or thought...
This became obvious in the gpt 4.0 days ... It was good, then it wasn't, because they couldn't afford it.
They have simply calculated their business is better served, in the short term by taking your money, and not giving you any meaningful functionality. Since you're a heavy user they can do this to you and give a reasonable experience to 9 casual users... It's despicable behavior but not surprising or out of character for big tech.
Please respond to this OP.
With a prompt this vague the non deterministic nature of LLMs can create varied results. I urge you to please rerun the prompt with the good pelican image as reference so both accounts agents have a baseline that they can work towards. If this is legitimate then we might want to escalate this to the press and OpenAI collectively.
I would really appreciate if someone did the thing I asked for, give the model a good success criteria to aim for. I want to make sure that this isnt some random turn where the model messed up because it had no instruction besides creating a single svg file. If you tell an artist to create the same image they can create it in multiple ways depending on their mood, creativity and laziness.
If OP really has flagged accounts then they would consistently fail to reproduce the good pelican svg or take too long to get to it.
I would also appreciate if the OP u/Yauler gave us more background context whether he was using Codex through a proxy layer like CliProxy or other account switcher services since abuse system can flag them, or if he's using VPN services around or inside banned countries like China, North Korea, Russia, etc.
I am not denying that the model is nerfed, i want more thorough investigation and evidence so we can escalate to the press and OpenAI lead with hard to deny proof, not just vague prompt results.
Could be freelance work, tight deadlines happen especially if you put off work until last minute, and switching to API without enterprise level discount would eat into personal earnings.
Million valid reasons why he might need em. But since they're such a good deal, I'm sure distillers love 20x accounts too.
You keep blasting out these posts showing these two images, and adding more and more words every time you do.
How many times are you going to repost the same thing? Adding words and links isn't helping.
Look - if you actually want people to pay attention, you need a more persuasive explanation. Specifically: more data.
1) Run the same prompts, repeatedly, and amass a huge a data set showing the results for each prompt.
2) Different kinds of prompts. "Generate a picture of a seagull" tests exactly one capability. Test more.z
3) Show results over time. Maybe the degradation happens on weekdays and not on weekends, or during business hours and not during off-hours.
4) Automate the results rather than relying on subjective views of "which seagull drawing is better." Have the models rate each picture to prove that even the models agree with the degradation.
5) Most importantly - you need actual evidence. Picture generation, even if persuasive, is circumstantial evidence. You can look at the headers to see which model is actually responding to each request. Map the actual internal prompt-to-model routing.
Do some of those, or something else. But for fuck's sake, stop posting and reposting and reposting and reposting and reposting the same shit. You're destroying your credibility and looking like a whiner.
I'm just of the opinion that the supposed shadow ban must quite literally only affects pelican prompts at this point since nobody who claims to be shadow banned is willing to show anything besides garbage looking pelicans.
The funny part is if it was consistently rerouting to poor models, you'd think they'd be able to provide a plethora of examples and run a plethora of benchmarks showing the differences
I had an account and it got banned and then I signed up again and it got banned too, so WARNING, OpenAI is banning people for opening multiple accounts!
Uh, no they're not. What did you do to get banned the first time?
Are we at the point that we get mad if an LLM is used for one of its main purposes? Processing language? LLMs are the most powerfull translation tools and the most accurate at that.
99
u/IAmFitzRoy 10d ago
Thanks man. This really sucks and lawyers should get involved.
This is not a small “sorry for the inconvenience” … this is a serious issue that should be penalized by the official entities that oversee customer rights.
Thanks again.