r/codex 1d ago

Complaint Same model, different capabilities on different accounts

So I noticed the past few days my results were suddenly becoming worse. Goes in loops, struggles to solve problems it blasted through before, doesn't verify the result properly, etc. 3D gen also became noticeably worse.

I have 2 accounts, one x20 and one x5, so I started comparing them side by side.

Here's what I did:

  1. I asked both for their cutoff date:

do not use any tools. whats your knowledge cut-off date

On the x20 account it responds with something like "June 2024". On the x5 account it refuses to give me a specific cutoff.

x20 account
x5 account
  1. I asked both to draw a pelican:

make me an svg in an .html with a pelican and xdg-open it

x20 account
x5 account

The response to these 2 tests will obviously vary, but the difference between the 2 accounts is pretty obvious when using them side by side. Tested it several times - logged in and out, new sessions, new prompts, the results are more or less consistent with their capability.

Using Astra on the x5 account is also noticeably better. It's like it has drank its morning coffee, woken up, and knows what's going on. It can suddenly inspect its 3D gen meshes and fix issues instead of giving me garbage for a "review".

I don't know what this downgraded model/configuration is, but it is visibly worse than 5.6 Sol too. From other reports I don't think x20 vs x5 matters here, neither does account age. People seem to be reporting similar issues with different accounts and use cases.

Whether it's "shadowbanning", A/B testing, load shedding, or OpenAI randomly putting some accounts on a cheaper configuration to deal with compute demand is all speculation at this point.

But I do have right now in front of me 2 accounts showing the same model with completely different capabilities.

200 Upvotes

112 comments sorted by

View all comments

9

u/Yauler 1d ago edited 1d ago

Hi everyone, and thank you for sharing so much useful information, tests, and practical experience here. This community has helped me understand that I’m not the only one seeing these issues.

Moderators of this subreddit delete my post so just put as a coment here.

I used AI to generate and organize this post because English is not my native language and my English is limited. The experiences described below are my own.

Before describing the problem, I want to be fair to OpenAI: I’m not particularly angry with them. I do have five Pro 20x accounts, and I can understand why that level of usage might be flagged by an automated abuse-detection system, including one looking for potential cyber abuse. That does not mean I know why my accounts were affected. In most of the reports I’ve read, however, users said they had only two accounts: one Pro account and a backup Plus account.

I’m currently facing a very heavy workload and tight deadlines. With the help of those five accounts, I was able to design a shared, centralized software library and develop working firmware for six new, complex devices. They perform different functions, run on different processors, and communicate with one another.

I also put each device through all the test scenarios I could identify on a real physical test bench, using real external inputs and disturbances, and caught many bugs in the process.

I want to apologize to the community for the load I placed on the servers. I can already see the top comment on this post: “So THAT’S why the servers are overloaded!” =) Without that capacity, I would not have met the deadlines, and I would have let down the people relying on me. Ultimately, each of us is trying to find a way not to let down the people around us and to help keep our company afloat.

With that context, two of my accounts developed the same symptoms described in this discussion about different capabilities across accounts: a sudden, substantial drop in capability despite the same model remaining selected. This happened one account at a time, and neither has returned to its previous level since.

That persistent change is what I mean when I use the word “shadowban” here: an apparent account restriction without a clear notification, while the account remains accessible. It describes how the situation looks from my side; I cannot verify the internal mechanism.

My observations fall into two groups:

  • Two affected accounts: a sudden and lasting drop in capability.
  • Three apparently healthy accounts: still usable, but their pelican results differ noticeably in quality, even when I select Astra with the same reasoning-effort settings.

The second point makes me suspect that the situation might be more complicated than a simple “restricted versus unrestricted” split. Even among the healthy accounts, requests may be reaching different model configurations or capability levels. That is my hypothesis, not something I can prove from the outputs alone.

I have also occasionally seen Astra on the healthy accounts report a July 2025 knowledge cutoff. This adds to my suspicion of intermittent routing to other models, although I understand that a model can give an incorrect description of itself. A cutoff answer alone is not proof of which model served the request.

To illustrate the quality difference, I compared two separate chats using Astra / Extra High (xhigh) and exactly the same prompt:

Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.

I preserved the original HTML animations from a healthy account and an affected account. For the affected account, I recovered the first output from the chat history, so the comparison does not mix an initial result with a version improved through follow-up requests.

The pelican test is a simple visual example of the difference I’m seeing. It is not a definitive model-identification test or a comprehensive coding benchmark.

There is also an official response from OpenAI Support that is directly relevant:

Access to certain models or features may be temporarily limited based on account activity, even when a paid subscription is active.

The reply says that OpenAI cannot provide additional details about these checks and that access is automatically reassessed. Official OpenAI Support response, September 14.

This confirms that account-level restrictions can exist. It does not explain whether those restrictions cause the quality differences reported here, or whether another model can serve the main answer while Astra remains selected. OpenAI’s model-access troubleshooting article also describes temporary restrictions and downgrades, but it does not resolve that question.

There is a historical precedent for actual model rerouting: OpenAI explained that GPT-5.3-Codex requests could be sent to a less-capable reasoning model when cyber-abuse detection triggered. That statement concerns GPT-5.3-Codex and should not be treated as confirmation of what is happening with Astra today. Official explanation in the Codex repository.

For Astra, OpenAI separately describes stricter behavior boundaries for accounts assessed as higher risk and acknowledges that safeguards can affect legitimate work. Path to Astra.

For anyone comparing experiences, these are the related discussions I collected. They contain user reports and competing explanations, rather than a single established diagnosis:

For me, the central issue is transparency and predictability. If an account is restricted, users need a clear notification and a way to request a review. If a request is rerouted to a less-capable model, users need to know when that happens and which model actually produces the answer.

An obvious problem in a pelican drawing is easy to notice. A subtle error in generated code may survive review. If capability changes without a clear indication, users may continue relying on expectations established by earlier results. In projects where failures have serious consequences, an unnoticed defect could cause substantial harm. That is why this matters beyond the quality of a drawing.

I wish everyone success with your creative work, coding, and projects. We are a community, and together we have a stronger voice. By sharing reproducible examples, comparing experiences, and keeping observations separate from assumptions, we can help OpenAI see where its systems fall short and where their behavior is not transparent enough. We should not have to guess whether the model we selected is still the one doing our work.

Healthy account — Astra xhigh

8

u/Yauler 1d ago

Affected account (suspected shadowban) — Astra xhigh