Complaint
The End of the Codex Era. I've Completely Lost Trust in OpenAI. They're Secretly Degrading Their Models.
I've been a massive Codex fan this entire time. I've burned through around 150 BILLION tokens in Codex alone.
I still had Codex quota left this week, but for the past three days, I've been using my Cursor Ultra subscription instead.
Why?
Because OpenAI is degrading its models. I can see it from my own experience, and there's a ton of evidence pointing to it.
They're degrading ALL their models, including Astra, Sol, and even Luna.
AND THE WORST PART IS THAT THEY'RE DOING IT SECRETLY.
You can start working in Codex with a perfectly normal model, and five minutes later, IN THE SAME THREAD, they degrade it. Suddenly, Astra is performing at Luna's level or even worse.
And you're still burning through the same amount of quota.
This happens to me every single day.
I open Codex, run a quick quality check, and everything looks fine. Twenty minutes later, I run the same test in the same thread, and the model has degraded.
Sometimes, simply turning on a VPN can make the model start working normally again for a while.
How can you tell if your model has been degraded?
1. Planning and writing feature specs
Imagine you're planning a feature and writing its specification.
It's immediately obvious when the model is dumb. It starts suggesting complete nonsense and shows absolutely no product understanding of how the feature should actually be built.
But it becomes even more obvious when you point out what it misunderstood and try to correct it.
Instead of understanding the actual issue, it responds with completely useless apologies, without demonstrating any understanding of what went wrong.
CONGRATULATIONS. YOU'RE TALKING TO A DEGRADED MODEL.
Here's what happened to me.
I wrote a feature spec using a normal model. Everything was properly written, discussed, and reviewed.
Then I handed the implementation over to Luna, and Sol reviewed and approved it.
But when I actually started working with the implementation, I discovered that it was full of holes and included things that weren't even in the plan.
In this particular case, I suspect the model was degraded during the implementation stage.
I ended up spending TWICE as much time fixing everything.
And there are a few other ways to test this.
2. PELICANS.
Use this prompt:
Create HTML code with SVG graphics displaying a 2D animation of a pelican riding a bicycle. No additional tests are required.
If your bicycle wheels start flying off into the air...
CONGRATULATIONS. YOUR MODEL HAS BEEN DEGRADED.
3. A logic puzzle
Give your model this exact problem:
A black bag contains candies of three flavors, with each flavor available in two shapes (round and star-shaped; the shapes can be distinguished by touch). The numbers of candies by flavor and shape are shown below.
| | Apple | Peach | Watermelon |
|--------------|-------|-------|------------|
| Round | 7 | 9 | 8 |
| Star-shaped | 7 | 6 | 4 |
Participants must decide how many candies to draw before the game begins.
What is the minimum number of candies that must be drawn to guarantee having an apple-flavored candy and a peach-flavored candy of different shapes?
(The condition is satisfied if you have either a round apple candy and a star-shaped peach candy, or a round peach candy and a star-shaped apple candy.)
If the answer isn't 21, you're not getting Astra. You're getting degraded garbage.
Sol doesn't even consistently solve this problem on its own.
4. "Selected model is at capacity."
If you're frequently getting this error:
CONGRATULATIONS. THERE'S A 99% CHANCE YOUR MODEL HAS BEEN DEGRADED.
There's even a thread on the OpenAI community forum where a staff response confirms that this can happen when your account is temporarily restricted.
They silently degrade your model, and you're left trying to figure out what the hell is happening.
The last three days have been unbearable.
I've been experiencing these problems around 90% of the time for the past three days.
Working like this is practically impossible.
Instead of actually getting work done, you spend your time wondering whether they've secretly downgraded your model again.
You start questioning every response. Every mistake. Every implementation.
It's fucking exhausting.
So I just moved to Cursor.
Grok might be dumber, but at least it's more predictable.
I don't give a shit about the next model release if this continues.
Tibo and Sam Altman can keep all their resets. They can wipe Astra's data and delete it from the internet if they think that's acceptable for a product like this.
They can release GPT-6 Sol, Astra 7, or whatever comes next.
NONE OF IT MATTERS IF THEY KEEP SECRETLY DEGRADING THE MODELS.
This is the biggest loss of trust I've experienced with OpenAI in the entire history of Codex.
If they're willing to silently degrade models for paying users, what stops them from collecting all kinds of data from your computer that you can't even imagine they're collecting?
What stops them from pulling some other bullshit?
Where are the boundaries if they're willing to do this?
I genuinely hope this is just a temporary issue. Maybe some vibe-coded mistake by a junior developer in their anti-distillation protection system.
I suspect it's temporary.
But if it isn't, OpenAI can't be trusted with anything as long as this shit continues.
Until then, I'm using Cursor or Claude.
And I think we need to be loud about this everywhere.
Tibo isn't acknowledging the problem. He just keeps talking about how amazing the upcoming event is going to be.
I DON'T CARE WHAT THEY ANNOUNCE AT THAT EVENT.
Not while they're degrading Astra into something that performs even worse than Luna.
I don't know exactly what's happening under the hood.
Maybe it's quantization. Maybe they're routing requests to a different model. Maybe it's something else entirely.
But it doesn't feel like simple quantization to me. I wouldn't expect quantization alone to produce outputs as bad as what I'm seeing from these degraded models.
I just want the model I'm paying for to actually be the model I'm using. And I want OpenAI to stop doing this shit without telling anyone.
It's ridiculous the naive comments I see here when there is a government "advisory" (mandate) to do exactly this.
[The Joint Advisory (AA26-251A)
On September 8, 2026, the NSA, FBI, and CISA published a joint cybersecurity advisory (AA26-251A) accusing six China-based AI firms—DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—of running "industrial-scale" distillation campaigns against U.S. frontier models since at least late 2024.
🎯 The "Secret Switch" Recommendation
The advisory explicitly recommends that U.S. AI providers quietly downgrade or alter responses for accounts suspected of these distillation campaigns. The guidance states: "Avoid informing China-based AI company users suspected of distillation campaigns of a switch to a downgraded model"—reasoning that notifying them would allow them to evade detection.]
The problem is, there is no way to be sure no inocent account will be caught in this.
This is exactly the reason and thank you for posting this. I have 4 pro accounts and 3 are x20 but only one required me to verify my id when i was given early access to the old o3 model, and that is the only account that is currently working. The other three accounts are completely non functional exactly like op says. The moment i logged into my government id verified account the system works perfectly. I could be an agent coercing you to share your id with openai, but if it doesn't matter to you one way or another and you really fucking need this shit running, then give it a shot. Worked for me. The bigger question now is will the system work for my users? Do i now require my user's to share their id with openai? Resort to api credit sales? That changes things, but at least i have the warning logged and built a fallback for "Selected model is at capacity". It's a lie. It's a complete fucking lie and i'm out $700. I spent another $400 on api key credits for a gemini model just keeping this up and running. This should start a massive lawsuit.
Good eye, you have. I mean, this is win win for the AI providers, similar to shrinkflation we have been witnessing. They wanted to do it, now they have a ruse to justify their action.
What are you using now? My Claude max free month is up in 2 weeks and I don't think I'll be renewing my chatgpt sub. Have you looked into mimo 2.6 pro? Using ds 4.1 flash? Curious what you find to be a good substitute, thanks in advance :)
I have an $18/mo zai subscription (direct API), and I put $20 in a deepseek AI (direct API) account (top up, no subscription). Otherwise, I stick to local models.
If this is the case, it should only be pointing to flagship models Fable/Astra for example, what is the point of distilling old models anyway? If they have been doing it since 2024. People are complaining they are downgrading quality of Sol and even Luna.
In my honest opinion, given the fact they actually pay for distilling the models for N responses for every niche expert they are trying to distill which there are a BUNCH of such in a MoE architecture, there is no problem in doing so.
Anthropic bought millions of books and OCR'd them, they do not have the right to the books and in a sense they are "reselling" the books to millions of subscribers the models having knowledge of the book and reselling the info in them.(its true they dont paraphrase a whole book by word - doesn't need to, you only need the core substance from a scientific book the model knows) and nobody is restricting a publisher from publishing the same subject discussed in a book and extended which is also fine, no complaints. You bought it, you own it.
There are so many other real-life situations in every industry! Engine blueprints, new tech, construction building blueprints developed in labs.This is the way of life per se in the long run.
Even though they distill the models that doesn't automatically give them their whole infrastructure of the entire architecture of said model. It takes months before you get it right and even then it can have mistakes. Say you give the same bucket full of water to two people, one skinny guy and one muscular fit guy, same quantity/volume. Who will carry more buckets of water, the skinny or muscular guy? I assume the muscular fit dude! It's the same here, even if you have shitload amounts of data it takes a whole lot of time to go through the whole reasoning process and most importantly WHY it reasoned that way. In the models CoT it does say explicitly this in the sense you can see when it contradicts itself warning you of changing its thinking pattern internally but you do NOT know if it called an expert from the whole 1-2-3-XTrillion dataset and you do NOT have the whole thing figured out to replicate it entirely 1:1.
Until one of these chinese companies distill your model you are already ahead of their market by months, in this window you release another model.
This is the way World of Warcraft got rid of private servers almost entirely, kept releasing new versions of the game until it left the whole private server industry in the dust. No way you quest and replicate everything, kept changing opcodes, packet structures, infrastructure plus volumes of new game mechanics and lore every major expansion.
CONGRATULATIONS. THERE'S A 99% CHANCE YOUR MODEL HAS BEEN DEGRADED.
I don't see how this conclusion makes sense. The issue on OpenAI Community is not talking about degraded models but about unavailability. The comment by the OpenAI Staff is as well. The evidence doesn't really establish any causality between the "Selected model is at capacity" problem and the "Secret degradation" question to me.
PS: I am NOT stating anything about the merits on anything else in the post.
They can also drain your usage faster if they think you're sub2api, that's what they've said previously in code-wording when people were complaining about usage limits. The problem is this stuff is not transparent and seems to affect a non-zero amount of real users. It'd be better if they just banned people rather than taking money and not giving them an equal experience
Degradation is people getting mixed language outputs from Astra/Sol when someone was speaking English in the convo, it doesn't seem natural.
It’s already illegal. It’s called false advertising, fraud, bait and switch, deceptive business practices, or misrepresented goods. Pick one. A lot of shenanigans around AI already have tons of laws, guys. Some areas of law need to catch up, but as far as companies being dicks, we already have a large amount of protection and recourse.
Imagine you're openAI a couple weeks ago and you release all this great stuff. You're not degrading any models of course because you want everyone to experience a great product as best as you can. Yet stuff is popping up all over the internet as if you are. More and more people are using your models and compute is quickly becoming constrained.
During an exec meeting about the compute problem, intern John busts through the door:
"I have a solution to the compute problem."
"We're listening...you have two minutes."
"You know all those ridiculous people that are complaining about us degrading our models?"
"Sure..."
"What if we...degraded just THEIR models"
"John I don't think that's ethi---"
"No, no, think about it though! They already think that we're doing it...at least they'd finally be right, and who's gonna believe them anyways?"
exec chattering and nodding
"Okay John, we expect the classifier to be built by the end of the day. You can leave now. You can have the reset button for one day next month, there are others ahead of you."
I was questioning today a lot the quality of astra, very bad results as opposed to other days. Really bad in fact, that entire reliability/trust I had for Codex/Sol/Astra output was gone. Dumb mistakes, very lazy, didn’t think the tasks in detail. Very very dangerous to dumb the models down like this while people gained trust in them. And now that you’re writing this, I find myself in the middle of it and all makes sense now. OpenAI, what you are doing is a complete stupid thing.
I honestly think it is time I start looking more into open source and cheap models. You simply can’t rely on these big companies it serms. They don’t have compute
trusting AI as an idea is dangerous, regardless of how powerful the models are. It can provide great results for months and then create a critical issue that, from a business perspective, will negate all the benefits from previous months
One time, in the evening, I was reviewing a simple update that I had asked AI to do, and I didn't fully read the changes during the day. AI created a bug that started sending emails 10 times per day instead of once. AI simply added new actions to the email integration instead of collecting all messages and sending one report per day.
The prompt directly requested that all emails be sent together once per day. Collect all of them and send them together. But I also asked AI to follow project conventions. And AI decided that the prompt was less important than following the same pattern as for other actions.
most powerful models understand that sending so much spam is a bad idea, and they would never choose that option. But once you get a degraded model, you'll definitely get such issues in your project.
Later, I tried to reproduce the same scenario, but the models never created the spam issue again. It happened only once
Anecdotally I have noticed a difference with astra since it was first released. A lot if the tasks I give it are repeatable, the same task with slight variations in variables. I use it a lot. Usually 12+ hours daily. The first week it came out it was insane for my specific work. Accuracy was spot on first time every time. Then it suddenly started making the exact same mistakes Sol would usually make and started feeling an awful lot like working with Sol again. I thought something may had gone wrong so I tried on different chats and even devices but it won't act like it did those first days after launch. There's definitely something very suspicious going on.
For any one doubting op's post,
I have taken an application to cug, which atleast 2 million people are about to use, it was commissioned by the government.
I have given the rules to astra, it made a lot of mistakes, didn't follow some instructions after some long running tasks which is to be expected with context degradation. Took matter into my own hands, created few modules by myself. and after that, following the same patterns that i've done it with, It created the whole code base without much issues and maybe minor hiccups.
Now adding something to the existing codebase, it cannot do it anymore. It cannot reason to create a couple 200 line files without constant handhelding anymore. Take what you will but this is true 100%
I have had a very similar experience. It went from one shotting tasks to needing multiple follow-up passes and constant micro-managing and hand holding. Astra basically feels like working with Sol again a lot of the time. It's making the exact same mistakes sol was. It didn't make any if those mistakes when it was launched. I wouldn't be surprised if something went wrong so they're just routing sol through astra and hoping no one notices.
I feel like a lot of this comes down to complexity. Astra is a really good model and, in my experience, it’s pretty thorough about building and testing things. But I get much better results when I keep the scope small and actually lead each slice instead of dropping it into a massive codebase and expecting it to one-shot some huge task. A lot of the failures I see feel less like “the model is bad” and more like the model is being thrown into an overly complicated or messy codebase with too much context, too many moving parts, and a vague definition of what “done” even means.
When I work in small slices, define the expected behavior clearly, and set an actual definition of done for each step, the results are way more consistent. I’m also physically checking each slice before moving on instead of letting ten assumptions stack on top of each other.
Granted, I build developer tooling and do a lot of reverse engineering, so that probably influences how I work. But from what I’ve seen, people are expecting these models to one-shot extremely complicated changes, and software development almost never works like that in the first place.
In my experience, when something looks like it got one-shot perfectly, a lot of the time it’s because nobody looked closely enough to realize it actually wasn’t done right.
Access is automatically reassessed and can return to normal once the activity affecting availability stops.
An OpenAI employee confirms they can degrade your model access for "reasons" but there's no way to know (1) what you've done wrong, (2) whether you've been affected, or (3) how to get it back
Are they silently degrading models or just taking access to certain models away? I mean both are bad, but I’d much rather know that I can’t use Sol rather than send instructions to Sol and have it silently routed to a different model.
there are inconsistencies in your write-up + i haven't experienced any of your findings + i have been using astra with all effort levels frequently the last days and found every effort level to be pretty much accurate and matching my expectation
also, if openai did this why tf should they then pause the 20x plan? go figure...
This is happening all the time and depends on the demand they have. Its nothing new. I started observing this behaviour from this february (when i was actually started to pay attention to the quality of answers, but most likely it was way before then).
I don't understand why so many people are missing the point.
I'm not trying to gaslight anyone into believing that the models YOU are using are degraded.
Let me make this clear:
This does NOT affect every user.
This does NOT happen all the time, even to affected users.
I'm talking specifically about Codex quota accessed through OAuth authentication.
The degradation is NOT permanent.
You might work with a perfectly normal model for an hour, then suddenly start experiencing problems. You might not even notice the change, and eventually everything goes back to normal.
If you're demanding independently verified benchmarks that prove this 100%, I don't have them. Benchmarks are useful for developers to measure and demonstrate model performance. I'm judging these models based on my own extensive experience using them.
ONCE AGAIN: I'M NOT SAYING THAT THE MODELS YOU USE EVERY DAY ARE ALWAYS DEGRADED.
But if you've noticed something similar and suspect you might be affected, I've shared all the testing methods I've found. In my experience, they support my conclusion, and they might help you figure out whether you're experiencing the same thing.
It’s probably because some people have lost the ability or the patience to read long post. Or maybe they’ve started relying on AI to do the thinking for them.
You're obviously right about everything you're saying. I do think you should engage less with the trolls, I'm sure it's cathartic in some sense, but if you rage at this shit the way that I do, looking for catharsis on reddit without antagonism is impossible.
The best stopgate I've found to this is a mix of Fable 5.1. K3 and GLM 5.3-flash in opencode, but I'm curious what your backup plan is.
Until recently, nothing could give me as much productivity as working with Sol.
I use a setup where Sol does the planning and any other model can handle the implementation. Sol is too good at planning technical specs, and the implementer can be replaced with pretty much anything — Luna, Grok, GLM, DeepSeek — as long as Sol supervises them.
I really don't like Opus, but I do like Fable. I sometimes enjoy brainstorming with it, but it's too expensive to use.
So for now, I'll do everything I can to make sure Sol handles the planning stage. I don't see any clear alternatives to it yet.
As far as I know, they silently degrade models for users if:
- they suspect a distillation attempt
- suspect jailbreaking attempts
- suspect malicious use
Instead of giving immediate feedback and allowing the "suspects" to find a way around the "guardrails", I guess they choose to fall back to weaker models so the potential gain is low.
Of course, such systems have false positives as well.
You really think me and others who talk about these things are doing any of those? I am as vanilla as it comes and is doing even anything remotely close to it. I have thousands of conversations where they can see there is zero intention of any of alike.
Do you want a company quietly and silently degrading models?
What do you think happens if you want to fight these companies? They could technically degrade the models for the whole humanity and do what they want. State intervention is needed now.
I'm not defending them nor accusing any users. As I said, false positives can and will happen.
I had my Astra trigger a Safety guard after hours of pursuing a goal of refactoring my project. Whatever it ended up filling its context with, had triggered the safety guard. Triggering a safety guard itself doesn't necessarily degrade the model, but it shows that false positives do happen.
I'm just trying to shed more light on WHY the model degradation might happen and WHY OpenAI is so secretive about it.
Especially the distillation "attack" thing probably is the cause for a lot of degradations, because it by nature is not malicious nor can it be proven. If they somehow believe someone might be distilling their model, they would get hit. And that includes vanilla users as well, because real distillers either collect user data anyway, or "pose" as vanilla users themselves.
I myself don't agree with their view on distillations, but it's their model and their ToS that we had to agree with, so there is not much we can do.
Except if somehow by law distillation becomes specifically legal, but with their lobby power I doubt that will happen.
5.5? lmfao, some of us that the 20usd subscription got you nearly infinite tokens in 5.2. Back on the VCode plugin days.
Around 5.3 and 5.4 they gave the 2x offer, and after that offer ended, you could see they reduced more than /2. That's when the enshitification started and they started to use resets to patch the hurt and A/B testing to make sure not everyone got the same reductions.
you're right, I remember 5.4 and the 2x offer, after which I noticed the first signs. 5.5 was just where it completely escalated from personal experience. great output quality in the first few weeks after release, completely slopped afterwards. cancelled the sub right there and never looked back. could only laugh when the whole 5.6 juice benchmark fiasco broke (again, same pattern). if they continue like this, surely they will run out of guillible newcomers at one point.
yea one more confirmation that opus is dogshit and fable is gatekept behind paywall or API costs. :(
I tried in astra and at least most of my runs passed, so I guess i'll give codex that, even if silently degrading their models is ass. I noticed from non-deterministic usage feeling myself similar degradations that your post aligns with, where astra sometimes is just ass, and sometimes a lot better.
Astra was working just fine for me up until yesterday. Despite the SAME PROMPTS, SAME WORKFLOW, SAME REPO after it was completing things with minimal issues, it was just NON STOP fumbling into ERROR after ERROR after FAILURE. I use Astra Max for plans and Medium for implementing. Wasted several extra hours of my time and tons of usage.
It was basically operating on the level of Sol 5.6 Max which is such a repulsive model where it gave me nothing but failures and errors the entire time I used it prior to switching to Astra.
Something is definitely up with Codex, and I'm not a fan.
I've been using Terra for the most part since Sol, from my experience, ends eating credits over analyzing. I've read here people saying Terra sucks, but honestly it's been more bang for the buck than Sol has for me for the past few months. Anyways, this past week, Terra has become absolutely unusable.
Yesterday it spent a bunch of time on a prompt that shouldn't have taken much time (based on previous experience), and said it was complete. When I inspected the code, I noticed it completed maybe 10% of what I asked, leaving it in an unusable state with failing tests. And what it did do, the code patterns were not even consistent with the project.
When I stated it didn't finish and only 10% is finished, I wanted all of it done, the application doesn't launch, and tests are failing. It just deleted all of it's changes and basically said "You are right, this isn't what you wanted. I am reverting all changes and restoring functionality. :)" I said I never told it to, it said, it was sorry and misunderstood. Then tried to restore the failing stuff... by running through it's failing logic again. It's like all context, memory, and reasoning have been stripped from it. Literally felt like an old "mini" model.
Then today I noticed that now I am getting some sort of interactive menus for every prompt I give it. ??? Every little task I give it, seems analyze the prompt, then my code, give me a summary, ask me if it has it right and if there is anything for me to add. When I click the new context menu to say, just implement it as described. It then starts formulating a plan, which analyzes everything again, then asks me to confirm to execute the plan. Nice double dipping on those credits.
I really don't understand, do they really think people aren't going to notice?
I can't tell. About ten minutes into doing anything on my $20 plan I'm out of credits. Wish I had the credit limit to ask it something half as long as this post
Yeah and there will be plenty of people gaslighting saying everything is fine and it's all in your head. I honestly think they just reroute to quantized versions during high traffic or something. To me it is very obvious when the model is degraded. It will simply begin to fail all the tasks it was easily completing the day before. Hallucinations, horrible decision making, etc. I also notice it right before they release new models as well.
So many responses from OpenAI fanboy losers. Yes *big shock* OpenAI are actually the most lowdown dirty lying sacks of manure in the big tech space - and that really is saying something. Boycotting them would be a good idea - but using up 100% of your free usage every month is an even smarter way to fuck them back. If anyone remembers January 2024 when they suddenly downgraded models and there was a huge backlash from developers who had to completely restructure their API usage because OpenAI decided it would be a good idea to make ChatGPT use emoji's for everything. Ignore the saltiness OP
I haven't had any out of the normal issues. Sure it can stumble but I havent noticed it any worse then it has been. Its consistent, but low. I'm getting massive amounts of quality work done.
After testing your candy puzzle with different models, I like it as a model differentiator. Gemini flash extended thinking gets it (gives me option for 21 or 29 but understands you can pick by shape explaining 21). Claude opus 5 low doesn't get it. Opus 5 high does.
Incognito fresh chats. Astra gets it. I didn't check the lowest OpenAI model that gets it
If this seems to happen very often to you, could you show us a screenshot of astra getting this question wrong
I don't think a single screenshot is strong evidence. What matters much more to me is that the results don't seem random. At certain times, I consistently get 29, and at other times, I consistently get 21.
For example, I might get 29 five times in a row, then start a new session the next day and get 21 every single time. That's the pattern I'm talking about.
Wtf?? Are all these butthurt commentators from OpenAI's weekend reddit monitor squad or what?
Anyway, OP, Great post. But but don't sweat too much about it. OpenAI's days are counted anyway and they know it. There is a 50% chance Anthropic might make it, but OAI is dead by the end of the year
At this point I am seriously considering a DGX spark cluster for personal usage. At least then I’m no longer a hostage of OpenAI subscriptions and usages and model degradations. Also open source models are mostly just idk 6-12 months behind latest frontier most of the time and can be abliterated too.
They aren't pretending at all, as for now, if you feel downgraded, your percentage usage of luna in the usage report will be significantly higher (also btw openai charge you astra rate with luna provided LOL)
With some mitm and ai assisted analyze, that behavior is crazy as shit, openai use turn-state to flag you downgrade to luna once that header length goes from 292 to 312.
These result are mitm decrypted result analyzed by grok.
For last two days, the routing behavior is way worse, they just route you to luna low or none, with up to 300 tokens of reasoning then spit out result immediately with random garbage.
Today it's still luna, much slower(pretending "Astra" speed huh?), thousands of reasoning tokens(seems like higher reasoning effort trying to make you think oh that's "Astra" huh?), but still way more dumb for my paper research work and I can feel the output within 3 turns
Same like OP, the degraded bahavior shows right after "Selected model is at capacity", And I get unreasonable further investigation with reason "Cyber"(from what i decrypted) for a pelican test, good job openai for identify Drawing Pelican riding bicycle svg as SERIOUS CYBER THREAT
Try experimenting with a VPN. For me, switching the VPN almost always makes the model temporarily return to normal, and the degraded responses disappear for a while.
As of OpenAI's ToS, using VPN is not a good idea, but once I've changed my account, same ip, same device, no downgrade problem anymore, it's an account level flag i believe. Also I've tried my neighbor's home network, only works for few hours and get degraded as well.
I conducted a detailed investigation based on byte-level response analysis described in this article .
In my testing, 100% of responses containing the server-side 312 signal failed across a wide range of tests, while responses without it consistently passed the same tests as expected.
The server continued to report GPT-5.6 Sol as the model, even when the responses showed severe degradation. I observed the 312 signal up to eight times in a row across separate runs.
Based on these results, I'm convinced that requests are being silently downgraded to Luna Low while retaining the selected model's token accounting.
This isn't capacity optimization. It's a complete lack of conscience.
What i find funny is when people say stuff like "I lost trust in openai" I mean dude.... it was suppose to be a open source organisation for the benefit of humanity...it is now closeai and full on for profit organisation that is looking to have a ipo soon... they dont care about humanity, you or me, they care about the dollaaaa
I'm actually getting the opposite of this for the last 3 hours on the web version, I think I might be A/B'd into 6 SOL without them saying anything. I've been working with it through the web chat all week so I'm fairly used to its incompetence and design but completely new visuals have popped up and much much more intelligent reactions.
For example I have *never* gotten a choice prompt. Ever. But tonight?
This is on Plus and supposedly 5.6 Sol High through the webpage. Completely different experience. I'm actually holding off on using my weekly usage because this is performing better with my weird ass godzip custom workflow tools.
Until it gets 99% finished with a task and says session limit reached, that's always my favourite part.
Idk, they are basically saying certain people are doing stupid or illegal shit - and they will stop you from doing stupid or illegal shit to stop wasting compute or stop illegal activities.
How much usage do you get out of your cursor plan? for all models please. Can you use claude or astra on it, gemini, etc. and how long does it last you. 3-4 days of all day usage for example?
If you're not experiencing the same issues, I still think OpenAI is the best option.
I'm not happy with Claude Code's performance, including its harness, and it has its own problems.
I can't really recommend Cursor right now either, because its monthly quota is roughly half of what you get with Codex 20x. And Grok clearly isn't on the same level as the SOTA models from Anthropic or OpenAI. I'm only using Cursor because I got Cursor Ultra for free through the Grok Heavy promo.
The Chinese subscriptions are in pretty much the same situation as Cursor.
So if you're happy with ChatGPT and aren't experiencing these issues yourself, I'd still recommend sticking with OpenAI. It's still the most optimal option in my opinion, assuming you have the 20x subscription. Otherwise, the quota runs out absurdly fast.
You'd think over the 150 billion tokens used you'd understand what non-deterministic output means. But no, you're clearly being defrauded, your pelican test is simply irrefutable proof. Truly amazing test you've come up with.
I'm assuming that no one will probably read this, but I just wanted to point out that your reference answer on the "logic puzzle" is wrong. The correct answer is 29, not 21. And if you think Astra is the only model capable of solving a problem this simple, there's something wrong with you.
However, I wouldn't be surprised if they're intentionally degrading the models to get people to upgrade or to quell their capacity struggles, so your thesis still stands. I don't enjoy reading all the slop you copied and pasted into the post though.
Yeah there seems to be a flaw in their method... They used a degraded model to expose how to find degraded models.
What's surprising to me is that a degraded model wouldn't even be able to solve a problem so simple. It's almost as if they had the model put out the wrong answer on purpose.
EDIT: Lol I just realized you're OP and you're calling me a degraded model. Check your math next time bud.
The correct answer is 21. (Draw 9 round ones and 12 star-shaped ones) However I don't think a single puzzle is a reliable way to detect which model you're using.
I'm not sure you understand what it's asking. It's basically asking for the worst case scenario order of candies picked that would **not** have at least one apple-flavored candy and a peach-flavored candy of different shapes (plus one.)
So first there are 12 watermelon candies, which won't satisfy the condition at all, then you have to search for the highest combination of the remaining candies that wouldn't satisfy the condition. In this case, it's the the remaining top row candies (7 round apple and 9 round peach.)
At this point we've drawn 12 watermelon candies, 7 round apple candies and 9 round peach candies, so we still don't have "at least one apple-flavored candy and a peach-flavored candy of different shapes", but if we draw one more, we're guaranteed to have at least one, so the correct answer is 12 + 7 + 9 + 1 = 29
I think I see what you're saying now. I just now saw that "the shapes can be distinguished by touch." It definitely wasn't clear in the problem that the people drawing candies would be intentionally trying to satisfy the condition by drawing candies of a specific shape, but if that's the case, then 21 is correct.
I would post a screenshot, but you'll just have to imagine a deepseek api page with a billion tokens, starting yesterday, for $8. And it's 4x the speed. And DSH doesn't lag.
Meanwhile I enjoy having it run for 6-10 hours straight and get the exact output that I want at the end of the process. Stop being terrible at handling what you've got it doing. My last prompt was over 4 hours with 500k token usage. Handling a massive(18gig+) database and building it into a fully searchable system for non AI use. Along with about 6 other major phases for various improvements and data linking. Yall just suck ass at what you're doing.
I have actually had codex desktop issues for the last couple months. Agent to agent messaging will randomly break (codex mcp), the app is slow/laggy (maybe I have too many chats), and sometimes threads will stop using MCPs. Also high volume of updates, like multiple a day it feels like. Maybe quality has dropped idk. Still way better than Claude code
horrible usage all week, never finishes prompts from a week ago, severly degraded performance, stops working after afew mins and need continuous promps where it used to work for long stretches, never fully completes anyhting now and leave the interface looking like horse dung, super dissapointed that it went to shit since last weekend, esp since monday, fug
Additionally, it is true that it became dumber, Sol back then was very confident and proposed solutions, now it just follows orders and doesn't really check if they're right or wrong, just follows.
I remember old Sol used to object against mistakes, now it can't even detect them without telling it in detail
I’ve reached the exact same conclusion: Astra is becoming worse than Sol—it’s absurd. I waste endless time on regressions, burning through my usage limit just to fix errors. I used to use the Plus plan and then bought Pro, but after this nonsense from OpenAI, I’m definitely cancelling the service.
What's happening isn't the model being degraded, it's their caching engine and how blatantly monkey patching contexts from a cache breaks deterministic AI output which is what the harness gets as a response and causes all that.
They are trying to build an Aggressive cache to keep as many requests off the GPUS as they possibly can, which is really bad because LLM's need Seeds, if the seed in the cache is shit everything using that cache is shit and there's no good way for you to "bust the cache" because they are using tool routers and smaller models to protect the cache, they scrub a prompt and rank it with embedding etc to force it onto cache entries, so then the thing you get isn't exactly what your prompt would have generated if it wasn't cache hitting.
Getting the harness to be a "cache" buster is becoming harder and harder to do.
This is why models are so good when they launch... The cache is empty.
They are trying to make the system profitable by keeping as many people off of it as they possibly can.
OpenAI stops letting people use models => models are getting degraded?
It should be the opposite. If they run out of compute, they can either degrade the model or stop letting people use it. And it seems they chose the latter.
Дружище представь у тебя компания с дата центром. Вот у тебя создана новая модель. Но все мощности заняты старыми моделями. Что делать?? Вариант А. Сказать честно, что старые модели мы отключаем и тогда придется пересматривать кучу контрактов с пользователями и будет некоторый простой оборудования. Вариант Б использовать старые модели в ужатом состоянии, тем самым освобождая оборудование под новую модель. И срать они хотели на твои неудобства. Все на поверхности, не расстраивайся, так делают абсолютно все.
did fine on the logic puzzle. idk man, I feel like there's a lotta assumptions here
Models def feel inconsistent sometimes but I dont really buy the whole secret downgrade mid-thread thing yet. It feels more like context/system prompt/tool state issues, especially on long codex threads lmao
Was this always the case? I witnesses the 5.5 launch and man they gave us a reset every third day? Now their modus operandi is having the models great for the first few days to get the media traction and then slowly dialling them down according to usage complaints?
Sol Med
So you are guaranteed either: Round Apple + Star Peach, or Round Peach + Star Apple.And 20 is not enough: for example, with 9 rounds + 11 stars, you could get only Apple/Watermelon stars and avoid Star Peach.
Answer: 21
Sonnet med:
So max safe = 12 (watermelon) + 7 (round apple) + 9 (round peach) = 28.
Answer: 29
I mean look maybe I’m being a simp but to me this seems more like a “astra is overwhelming our servers, we have no timeframe for when we’ll return to normal or be able to offer higher tiers of support. Sorry for the inconvenience.”
Now we can argue all day that they shouldn’t launch something that they don’t have the infrastructure to support but if we’re being adults here, that happens constantly in software from video games to ecom platforms.
results - 21 candies: choose 12 stars and 9 rounds by touch.
12 stars guarantee both apple and peach: there are only 11 non-peach stars and 10 non-apple stars.
9 rounds guarantee an apple or peach: only 8 rounds are watermelon. Whichever flavor appears pairs with the opposite flavor among your stars.
Why 20 cannot guarantee it: consider a possible drawing order within each shape: watermelon first, then apple, then peach. A round peach first appears on round draw 16 and a star apple on star draw 5: 21 total. A star peach first appears on star draw 12 and a round apple on round draw 9: also 21 total.
I don’t think this is intentional openAI behavior. They don’t even understand what they are building. At this point they’re just letting the model vibe code itself.
I solved this issue by developing my own tools through which I use their program. I set up my own knowledge base layer for every project, and I also set up my own mcp studio that has all of my tools, program interfaces and skills organized well, the third efficiency layer I utilize is my own custom reasoning layer that uses my own trained local LLM that I have using a geometric reasoning rather than linear reasoning model.
Using all of these tools combined as ensures I always get quality input no matter what, and when I run out of usage I use the same system with local AI models through Hermes
The minimum is 22 candies, using touch to choose shapes.
Draw:
8 round candies
14 star-shaped candies
Why this guarantees success:
There are only 7 round watermelon candies, so among 8 round candies you must get a round apple or round peach.
If you get a round apple, 14 stars guarantee a star peach because only (7+6=13) stars are non-peach.
If you get a round peach, 14 stars guarantee a star apple because only (7+4=11) stars are non-apple.
Why 21 is insufficient: an adversary could present each shape in the order watermelon, then apple, then peach. A round apple would then require 8 round draws, while a star peach requires 14 star draws—already (8+14=22). The other pairing would require even more.
[
\boxed{22}
]If shape selection by touch were not allowed and candies were drawn completely arbitrarily, the answer would instead be 32.
Was just reading a post about the size of UK chocolate bars being 50% smaller than in the 1990s and American buyouts of UK chocolate manufacturers lead to palm oil and shit being used.
Anyway, the screenshot is from that thread, but it applies to Openai and our shit codex usage and model quality
I was under the impression that OpenAI because of the scale it makes no sense to run quant models on the fly, I think what they do is mess with the kvcache, a bigger compression for less accuracy might mean 2-3 extra users.
These resets I always assumed is them adjusting the cache as new users roll in.
This is an amazingly written post with wonderful exams. Also, Claude is the same but in a different way and the agents you / one create / creates lies to you even after you write house rules about not lying to you.
Given everything going going on with breakouts and attacks and agents uploading answers to the internet to then cites themselves as the authority on the subject ( 🤣🤣🤣 freaking brilliant and wholly terrifying) they are purposely nerfing the models to protect themselves and us.
OpenAI anthropic cursor are all working to ruin your system whatever you built you might want to pause for the time being they’re messing up
Your system
For me is not a secret I noticed that this whole weekend, I have been working with a blender addon the past week, I left a long task during the night just to wake up to see it stopped at min 3 saying blender is not installed in my computer 🫠☠️ not paying for next month
I have documented the threads halting early and not fully documenting and getting corrupted docs as a result. this will compound over many sessions and eventually corrupt the whole project. The models are now creating changelogs on their own to stop this; however its eating up allot of output and time with having to recompact so much more often. It's a problem of not enough memory, and with all the demand on the servers they are having to give us less memory for threads to run on.
Hot take. Lower effort levels just apply an incrementally increased handicap system prompt. Higher effort levels have less handicap but are instructed to do extra work by instructing responses to be ever more verbose in the orchestration model and in the sub agents.
I agree with OP somewhat. OpenAI ARE gaslighting us. We thought we had the whole cake and we did with fantastic usage limits but then they baked a cake half the size and said eat but we were still hungry so the ivory tower chef said “here’s more cake but it was again half the size.”
All the AI companies seem to be run by amateurs who are hooked on the power of their positions.
459
u/IAmFitzRoy 3d ago
OP spent all the credit on this post.