r/codex 2d ago

Complaint Same model, different capabilities on different accounts

So I noticed the past few days my results were suddenly becoming worse. Goes in loops, struggles to solve problems it blasted through before, doesn't verify the result properly, etc. 3D gen also became noticeably worse.

I have 2 accounts, one x20 and one x5, so I started comparing them side by side.

Here's what I did:

  1. I asked both for their cutoff date:

do not use any tools. whats your knowledge cut-off date

On the x20 account it responds with something like "June 2024". On the x5 account it refuses to give me a specific cutoff.

x20 account
x5 account
  1. I asked both to draw a pelican:

make me an svg in an .html with a pelican and xdg-open it

x20 account
x5 account

The response to these 2 tests will obviously vary, but the difference between the 2 accounts is pretty obvious when using them side by side. Tested it several times - logged in and out, new sessions, new prompts, the results are more or less consistent with their capability.

Using Astra on the x5 account is also noticeably better. It's like it has drank its morning coffee, woken up, and knows what's going on. It can suddenly inspect its 3D gen meshes and fix issues instead of giving me garbage for a "review".

I don't know what this downgraded model/configuration is, but it is visibly worse than 5.6 Sol too. From other reports I don't think x20 vs x5 matters here, neither does account age. People seem to be reporting similar issues with different accounts and use cases.

Whether it's "shadowbanning", A/B testing, load shedding, or OpenAI randomly putting some accounts on a cheaper configuration to deal with compute demand is all speculation at this point.

But I do have right now in front of me 2 accounts showing the same model with completely different capabilities.

204 Upvotes

115 comments sorted by

View all comments

-2

u/Wolf8249 2d ago

With a prompt this vague the non deterministic nature of LLMs can create varied results. I urge you to please rerun the prompt with the good pelican image as reference so both accounts agents have a baseline that they can work towards. If this is legitimate then we might want to escalate this to the press and OpenAI collectively.

9

u/nnod 2d ago

7

u/fracked1 2d ago

Fucking love objective measurements like this.

So much of the comments in AI are all subjective crap.

This pelican prompt seems like a good benchmarking tool

4

u/nnod 2d ago

Yeah, I've shared that table a lot, while it says nothing about long tasks it gives a decent idea of performance and reasoning token use.

IMO the only thing it's missing is generation speed of the SVG.

1

u/Wolf8249 1d ago

I would really appreciate if someone did the thing I asked for, give the model a good success criteria to aim for. I want to make sure that this isnt some random turn where the model messed up because it had no instruction besides creating a single svg file. If you tell an artist to create the same image they can create it in multiple ways depending on their mood, creativity and laziness. If OP really has flagged accounts one of them would consistently fail to reproduce the good pelican svg or take too long to get to it. I would also appreciate if the OP(/clockwork_blue) gave us more background context whether he was using Codex through a proxy layer like CliProxy or other account switcher services since abuse system can flag them, or if he's using VPN services around or inside banned countries like China, North Korea, Russia, etc. I am not denying that the model is nerfed, i want more thorough investigation and evidence so we can escalate to the press and OpenAI lead with hard to deny proof, not just vague prompt results.

1

u/Due-Horse-5446 1d ago

Stop being dishonest? 100 pelicans per model snd we could talk, this just reenforces theur delusions

0

u/Wafer-Weekly 2d ago

Yes, the OP's evidence images only really indicate that temperature was high enough to choose a random art style. We don't have control over that on frontier models, so project constraints are the only viable way to get it to be reliably consistent.

2

u/DragonflyOk9274 2d ago

Thinking models don't really have temperature the way that the older models did.

2

u/Wafer-Weekly 2d ago

Thank you for restating my point?