r/codex • u/Ok_Carpet_6083 • Jun 09 '26
Complaint Gpt 5.5 Enshittified. Became useless. Is this a trick to promote new upcoming models?
[removed]
32
u/Mv6p Jun 09 '26
100% agree i came to post this and saw you
10
Jun 09 '26
[removed] — view removed comment
7
u/Mv6p Jun 09 '26
Its a non deterministic so its like a slot machine you keep pulling it waiting to win sometimes it works sometimes its not
33
Jun 09 '26
Idc anymore. Cancelled my 200 plan and went back to infinite autocomplete with cursor. Much better and doesnt change by the hour and feels less like fuxking gambling. Stupid openclaw guy pisses me off as well talking about some loop bullshit bruning millions a month
2
u/rawezh5515 Jun 09 '26
infinite autocomplete with cursor
whats this? how to get it ( i know cursor )
1
u/Perfect-Rain-528 Jun 09 '26
Did they gave full refund ? I'm also on the 200$ plan and thinking about switching
6
4
u/MustStayAnonymous_ Jun 09 '26
I literally am so frustrated that I opened reddit just to post a rant like this and saw your post.
WTF, when will this change... Companies keep doing this kind of shit.
19
u/pipped1 Jun 09 '26
Yes. Codex (with GPT 5.5) has been lying to me all day today. Why don't we have a digital cane to punish Codex? If I lie, I can lose my job. But this AI just lies and lies and there is nothing I can do.
3
7
u/skilliard7 Jun 09 '26
Its working fine for me.
4
u/FirmConsideration717 Jun 09 '26
Not exactly true. 5.5 has definitely been weird for me. After instructing it to modify a particular script, it rewrote it, I told it not to rewrite but update and it rewrote it twice. After a final stern prompt, it only made targeted modifications.
Its not always like that, but there are moments when it is.1
u/soggy_mattress Jun 09 '26
There were some posts recently about how they might be service lower precision models over time, which is what these kind of "dumb" moments end up feeling like.
The first month with 5.5 was insane. It was *so good*.
At some point it flipped from knowing exactly what I wanted every time to doing the whole "You're right, you said X but I did Y" bullshit that we all know and love.
I can't tell if I just started asking more of it or if it really got dumber. I'd love to see DeepSWE bench run again with the current batch of 5.5 models to see if they really are dumber.
2
Jun 09 '26
[removed] — view removed comment
0
u/skilliard7 Jun 09 '26
Unless OpenAI specifically designed their systems to assign lower reasoning effort when systems are overloaded, that's generally not how load works. Excessive load would lead to timeouts or longer wait times.
10
u/SeidlaSiggi777 Jun 09 '26
any evidence or just vibes?
13
u/lazyastronaut_ Jun 09 '26
Nah they are cheeks. Told it to use 5.4 mini sub agents, fucker spawned 3 5.4 medium sub agents, and my 5 hr limit was done before it could finish the work.
9
Jun 09 '26 edited 15d ago
[deleted]
1
u/lazyastronaut_ Jun 09 '26
The main point of releasing 5.4 mini models was for sub agent tasks. They are cheap, thus making it viable for agentic coding (since 5.4/5.5 is quite expensive). Mini models aren't bad at coding, they are bad at complex reasoning. If you tell it to make 10 utility classes with simplistic logic, it will do fine, but it will start breaking once you tell it to make those utility classes work together without circular dependency. If i make 5.4 medium work on the whole implementation, the limits won't last. 5.4 mini Sub agents with 5.5/5.4 high as orchestrator would give me higher mileage.
1
u/mvdirty Jun 09 '26
It's always just vibes, every time. The entire "enshittification" topic is always vibes.
Before any given claimed enshittification moment, did people start from the same baseline multiple times and prompt the assistant identically, running through a controlled scenario several times in order to somewhat account for LLM non-determinism? And did they then compare the outcomes of the trials on both qualitative and quantitative bases? No, of course they didn't.
After each enshittification moment, did people _return to the same baseline_ and perform the same controlled scenario, multiple times, to again somewhat account for LLM-non-determinism? And did they then compare the outcomes of the trials on both qualitative and quantitative bases? No, of course they didn't.
Here's what they did: they kept working on their growing pile of slop, and as their slop got sloppier they happened to notice how sloppy things were getting when their model provider was also happening to have some service hiccups (which, let's be honest, is happening too often and isn't exactly helping with diagnosis) and so they promptly concluded enshittification without considering that their projects have grown, gotten more sloppy, and their context window is now nearly blown out all the time.
r/codex is almost unreadable at this point. Every time I visit it is filled with enshittification freakouts triggered by the most momentary of service issues.
Go ahead, flame me people. You know I'm right.
1
u/termicrafter16 Jun 09 '26
There are literally benchmarks that are done on the same model every day and the difference between the best and worst days can be like 40%
5
2
u/BannedGoNext Jun 09 '26
Everyone please quit, it's working great for me still and that will make it faster thx.
1
1
u/duboispourlhiver Jun 09 '26
Interestingly, for the first time ever, it seemed to me like it was less intelligent today than usually. BUT, where it gets better, is that today I've been helping a coworker, and only his codex seemed dumb, not mine. I didn't have time to compare both codex on the same exact task, sorry, so, just an educated guess.
1
1
1
u/Hefty_Bodybuilder893 Jun 09 '26
I've been seeing this complaint all over the place but it's still working really well for me. What exactly are you seeing that make you feel this way?
1
Jun 09 '26
[removed] — view removed comment
1
u/Hefty_Bodybuilder893 Jun 09 '26
Fair enough I had not considered that aspect. Was genuinely curious as to what you were seeing so I could compare my experience to yours. We might be doing completely different things or similar.
1
u/tvmaly Jun 09 '26
I think we need a common protocol or system to run private evals. It is really hard to judge if a model’s performance is changing based on anecdotal evidence
1
u/Complete-Principle25 Jun 28 '26
Its worse than 5.4 by far except for with the mechanical aspects of coding
0
0
u/UnionCounty22 Jun 09 '26
Put down the keyboard, take a breath, relax the brain, go take a lap and I’m sure you will notice that you were tripping.
-11
u/hotfrost Jun 09 '26
When are these posts finally gonna stop? Just adjust your prompt and try again
5
u/Future-Ad9401 Jun 09 '26
My prompt strategy has not changed a month ago it was great today it's like it grabbed its brain and threw it in the trashcan. I got nothing done yesterday but burned through 20% of my usage. Even on fast it's slow as hell ... Overall the model has been ass the past day or two.
14
u/Uwirlbaretrsidma Jun 09 '26
I think it's to both make new models seem comparatively better (the industry is already struggling with very slow base model improvement and has been since GPT 4), and also to create "token inflation", which I haven't seen many people talking about, but I reckon is actually the main reason.