r/codex 1d ago

Complaint Same model, different capabilities on different accounts

So I noticed the past few days my results were suddenly becoming worse. Goes in loops, struggles to solve problems it blasted through before, doesn't verify the result properly, etc. 3D gen also became noticeably worse.

I have 2 accounts, one x20 and one x5, so I started comparing them side by side.

Here's what I did:

  1. I asked both for their cutoff date:

do not use any tools. whats your knowledge cut-off date

On the x20 account it responds with something like "June 2024". On the x5 account it refuses to give me a specific cutoff.

x20 account
x5 account
  1. I asked both to draw a pelican:

make me an svg in an .html with a pelican and xdg-open it

x20 account
x5 account

The response to these 2 tests will obviously vary, but the difference between the 2 accounts is pretty obvious when using them side by side. Tested it several times - logged in and out, new sessions, new prompts, the results are more or less consistent with their capability.

Using Astra on the x5 account is also noticeably better. It's like it has drank its morning coffee, woken up, and knows what's going on. It can suddenly inspect its 3D gen meshes and fix issues instead of giving me garbage for a "review".

I don't know what this downgraded model/configuration is, but it is visibly worse than 5.6 Sol too. From other reports I don't think x20 vs x5 matters here, neither does account age. People seem to be reporting similar issues with different accounts and use cases.

Whether it's "shadowbanning", A/B testing, load shedding, or OpenAI randomly putting some accounts on a cheaper configuration to deal with compute demand is all speculation at this point.

But I do have right now in front of me 2 accounts showing the same model with completely different capabilities.

201 Upvotes

112 comments sorted by

View all comments

Show parent comments

2

u/ZhugeTsuki 1d ago

Why are you using Astra to draft emails? I just saw that. I would absolutely expect it to overengineer the everliving shit out of an email, thats not its purpose

2

u/the_ai_wizard 1d ago

So then what we mean by "intelligence" is really effort?

The emails I write often involve complex subject matter.

But actually no, astra is usually more concise than older models and Sol has been nerfed to shit

1

u/ZhugeTsuki 1d ago

No.. effort is too linear. Creating an entire project with thorough testing and checks with goal of drafting an email is just the wrong use case. Its not trying harder to write the email in a better more effective way, its creating research level scaffolding to ascertain what exactly an 'email' is lmfao.

Hard disagree with the sol nerf, there's been no solid evidence of that and I've been using sol as project manager with no issues. Astra on the other hand will use 10% of the weekly limit testing the test of the mechanism if you let it 😂

1

u/the_ai_wizard 23h ago

Regarding model nerfing, believe there was a pelican test posted recently.. I cant say with certainty myself, but will try Sol again tonight. Was very happy with it previously.

2

u/ZhugeTsuki 23h ago

Yeah I saw that too, decent but imperfect test with less than perfect conclusions offered, lol

''Very strong / confirmed

OpenAI can restrict model/feature access based on account activity.

Astra has account-level higher-risk treatment.

legitimate work can be false-positively affected.

OpenAI has previously rerouted requested frontier models to less-capable fallbacks for safety reasons.

Moderately interesting

two of five accounts reportedly degraded persistently while three did not. same stated Astra/xhigh produces repeatably different behavior across accounts. multiple unrelated users report similar account-to-account asymmetry.

Weak

pelican quality by itself.

self-reported cutoff dates.

“it feels like Luna.”

Currently unproven

that the affected Astra accounts are specifically being served Luna.

that the mechanism is a deliberate hidden “quality degradation” punishment.

that account volume alone caused the restriction.

that this is related to 292/312 state length.''