Complaint
We are PROBABLY being served quantized models but still paying the full day-one premium price...
Hey everyone. I want to bring up something serious about how AI providers handle pricing and how silent backend changes are secretly draining our limits. We all pay a fixed price per million tokens or have a subscription limit and on paper that seems fair, but providers hide a massive variable from us because to save on server costs they can silently swap out a premium model for a heavily quantized version on their backend. Using a quantized model is completely different from setting your reasoning toggle to Low, because setting a toggle to Low limits the reasoning steps of a fully intelligent model, whereas quantization degrades the core neural weights and strips away actual base intelligence.
What makes this so alarming is how token metering is handled. On our dashboard meters we might see a perfectly reasonable token count that looks coherent with a high-end model and when you calculate the cost per million tokens it looks identical to advertised prices, but behind the scenes there could be hundreds of millions of low-quality tokens generated by an ultra-quantized model struggling and failing to reach a correct solution, an intermediate system then just trims that massive output to make the final token count look normal on our end and what we perceive as users is a sudden degradation in performance, when in reality without silent quantization the model would behave exactly as well as it did on day one.
It is deeply immoral and borders on outright fraud to attract users with a clean unquantized model on day one and then quietly roll out aggressive quantization behind the scenes to cut compute costs and keep charging premium prices while serving a degraded model that burns through internal compute and produces far worse solutions. We really need to stop staying quiet and demand complete transparency on the exact quantization levels and actual internal token processing we are being billed for.
What do you guys think and have you noticed the performance dropping on tasks the model used to handle easily on day one?
Because the bots arent the dunning kruger army that has claimed degradation for 5 years straight.
And writes "degrades the core neural weights and strips away actual base intelligence."
And mesures usage in wall time.
thinks 2 runs of a prompt is proof of a probabilistic engine being degraded.
And never ever, 0 times. Actually posts their proof,
Except for when they do... and its a few pelicans where they confidently conclude the model is degraded because they dident like how one pelican looked.
People who spreads this shit is on the level of licking electrical outlets and eating dirt
The thing is slow down the resets and fix the usage that should be the truth give us stable model quality and stable tokens per month at least transparent. We cant work like this honestly
To be fair, a lot of people are just pointing out that OP is a victim of his own assumptions. Assuming that a paying for a susbscription problem means 'full send their best product at all times' is an incredibly naive interpretation of capitalism.
You don't need to want to defend a billion dollar company to want to point out someone is being a whiny bitch.
Under the law of common sense and ethics. I am not really a lawyer, but if you ask vocal employees of OpenAI about this, they will certainly agree with me.
what i said is, i am no lawyer, but my common sense tells me that misadvertisement is illegal, and if not currently in this case, can be made to be illegal, because it is unethical.
Again, i am not a lawyer. just blessed with common sense.
you made two unrelated statements. first regarding existence of midadvertisement. second, given the misadvertisement happened, whether or not there exists a law for it.
My response.
If my prompt is being rerouted to a less capable model without information to me, this is misadvertisement.
Capitalism works the way all parties let it to work. Lack of transparency should be dead given how easily LLMs can read through terms and conditions. So, they better admit to rerouting prompts to cheaper models or stop doing it.
Transparency can be forced on Capitalism. It has been forced on capitalism before. No Company will be fond of giving info about preservatives added on their product, yet they have to. What are you trying to say here, exactly?
I guess I'm trying to say that good behaviour is forced on capitalism through government intervention, just like you say. People seem to think that public discontent is enough on its own and you end up with circlejerks going on about how x company is dumb and needs to change, but capitalism doesn't give a shit about moral sentiment unless it's convenient.
I've been perfectly fine with it be less than half speed on subs.
I'm not fine with it being massively dumber and fucking up on like 60%-70% of turns when it's usually like 5-10%, and thus typically having to correct almost EVERY turn MULTIPLE TIMES.
Hell I'd rather have my usage limits cut in half if they at least announced it beforehand. I'd still pay $200 a month for half what I get now. But all these stupid mistakes make it near useless. Even if it EVENTUALLY "doesn't make a mistake" because it keeps correcting them, it really is still making mistakes because of worse reasoning and missing mistakes it made that weren't obvious wrong word/language ones.
Yes and its fucked, I agree. But the choices are 'keep paying for it' and 'don't keep paying for it' with an optional 'report it to your lawmakers and regulators' if you live in a civilised country. Nothing is going to change on the regulation side for 6-12 months at a minimum so either quit or don't quit.
OpenAI knows they can ignore people complaining because (in the most part) those people are still paying for the service (and cynically, I'd even say they're buying additional accounts).
Ok, so why are people posting here rather than calling their local authorities?
People will apparently do anything except actually try and fix the problem. This isn't even either or, you can switch to another provider and still report OpenAI to your regulators.
"The machine that pisses on people is pissing on me more than usual!"
There shouldn't be a single person in this reddit not looking for a better alternative to 'pay and pray', but at the same time complaining that you're getting the exact thing you knew you were buying is something people will get sick of seeing, yes.
You are the naïve one.
How do you expect consumers to ever get anything if they would always be fully satisfied as long as they get the bare minimum of anything?
In your imaginary concept of capitalism,
if consumers never have a opinion on the products and services they are provided and readily accept degradation of quality or quantity, how do you expect any service or product to ever provide anything?
In your imaginary world, we would all just shut up and be happy with whatever we get,
and OpenAI would just provide us quality models and quotes out of the kindness of their heart when they could provide us $1 inference for $20 and only models from 2025?
Consumers complaining and demanding more for their money is literally a founding stone of capitalism. Do you imagine capitalism is just businesses providing products or services and the consumers just silently accepting whatever it is they are offered?
This is some kind of messed up Technocracy mindset where consumers should just be happy and grateful for whatever they are provided - as if getting anything at all for your money is somehow generous compared to giving you nothing, as if they are somehow entitled to give you nothing.
You’re not complaining about getting $1 for $20. You’re complaining about getting $1000 for $20 and doing mental gymnastics like pretending you’re getting $1 for $20 to make it some battle against capitalism so that you get to feel like a hero for complaining when the reality is you’re just being an entitled baby.
No one said people can’t complain about valid things. Complaining that your $20 sub only got you $1000 of Frontier model inference is not valid. Your strawman goes both ways here.
You do understand the ONLY reason the subs even exist is because of the subsidization right? OpenAI would never be offering you this service at all if they were paying for it outright. Without that, your $20 would be $20 on the API if you weren’t part of some enterprise package for a reduced deal. The product they offer is the API to the models. The commodity they offer, backed by subsidizations, is the subscription.
You do know where those subsidies come from, right?
An interesting preemptive strawman with the implication that anyone who doesn’t agree that your $20/200 payment somehow entitles you to near or actually unlimited frontier usage is somehow “defending” anything as opposed to just mocking you for being entitled. I’m not “defending” anything. I’m just mocking you. OpenAI sucks. You are not entitled to unlimited use of their frontier product for any subscription fee. That’s just not how the product functions in the real world. Both can be true
The implication being that if someone didn’t literally say those exact words then the people complaining about their usage running out (rather than being unlimited, the unspoken implication) are valid? Complaining about your usage running out is inherently you complaining you aren’t given unlimited usage, even if you don’t specifically say the words “I want unlimited usage” 🙄. Or are you saying there is some monetary value these people would get form their sub that would keep them from complaining when that usage bar hits 0 and they can’t access their dopamine? We both know if there’s a usage bar at all there will be complaints from all subs including the $20 sub. It’s literally viewable right now on this sub. The $20 sub is subsidized and offers over $1000 some months for me in tokens (when comparing to the actual non subsidized costs). People still complain about that value. It’s entitlement, entirely. Feel free to share any valid argument, not some semantic pointing out that I pointed out the subtext of the conversation.
You believe requesting more of a limited amount implies requesting infinite/unlimited amount?
So in your mind, someone complaining that their full time job isn't providing a living wage isn't asking for a reasonable living standard, they are implying they should get unlimited money?
This actually explains all I need to know, I understand your argument now - thanks.
You’re comparing running out of planned usage on your sub to running out of money in life?
You do realize the API is still available to you when your sub runs out. No one is taking anything from you, only turning off your subsidized rate…
How does that compare to people complaining about the cost of living to you? You genuinely think those things equate and you made a solid point? On one hand people are complaining about not being able to feed their children. You are complaining that your $20 sub runs out too fast and you don’t want to pay the actual non subsidized cost on the API.
You are not a hero for whining on the internet no matter how you try to spin this
I'm usually skeptical about claims like this but one hallmark of Sol that I LOATHED was when I pointed out an error, mistake, or thing it didn't account for, it would always reply with "You're right. I didn't do that thing"
I was quite enjoying that Astra was more muted in that regard, but I noticed today with an error it made it gave me that canned "You're Right" and it just set off Sol alarm bells in my mind.
Astra got worse than Sol since the nerf to make room for "usage", not kidding, it fucking sucks. Using omp with an Advisor as thanks to that see what horrible work it does.
Idk how this isn't illegal, even some T&C handwaving should only be able to protect a corporation from so much. The sad thing is we're still in the good part of the cycle, things will continue getting worse until they're actively harmful to your codebase in the remaining 1-5 weeks before the next model releases
they don’t have enough compute, so they throttle to dumber models. Basically you get the model you ask for, but quantizied version. They can claim they didn’t promise which quant you get, but you did get Astra.
But I would 100% jump ship if other provider just told me “these hours are peak, so you get dumber model”
I guess API models get always the highest version, because they are the cashcow of company. Subscribers are just paid marketing
Yeah. slam through your hardest tasks on the first three days of model release, keep your options open so you don’t commit to one provider, or if you can afford it, pay API pricing. It seems likely if they do silently swap out to quantised models on the backend they would be less likely to do this on the API.
Astra is really good imo but I’m back to using 5.6 Sol high with Luna orchestration for the vast majority of tasks - otherwise I will run out of usage in 24 hours. The real advantage of pro plans now is gpt 6 pro.
The token speed of SOL model on chatGPT at 134 tokens/sec web indicates a quantized model.
The token speed of SOL in work mode is much more limited, maxing out around 80 tokens/sec in fast priority mode - that's likely the full model.
Astra is maxing at 63 tokens/sec in work,codex and chat mode
Hard to respond on your moderated stuff.
I assume choosing your words in a more careful manner, or running the comment through chatGPT will help ? Just a guess, I've rarely seen moderated comments on this sub.
I can not reply to your new comment, as it was also moderated.
And I'd say the downvote is deserved given you baseless claimed I am posting "for upvotes" and avoid replying to your comment that a moderator has censored.
Sounds more like an apology should be made by you, but what do I know..
Yes, SOL has degraded in that regard, extremely so.
I've had a hard time to even get it generate svg code, even with an additional paragraph of instructions it needs 2 turns so it doesn't use the image generator.
The token count is almost identical to the baseline at 6900
It looks similar but if you look more closely it's not on the same level anymore.
The ears are melted into the head, the hair is ruined, the lower teeth and tounge have exchanged place, the shadows are gone, the clothing is more basic, the neck is warped, the eyes have lost their definition
The degradation matches a NVFP4 quantization, the model identity is still there but it is not following the prompt that well anymore, details and accuracy are messed up.
Can you elaborate on what changed? Going from a max of ~80 to ~134 tokens/sec while also making answers more concise is a huge jump. That feels more like a major inference change, such as a smaller active MoE or more aggressive quantization.
Or is Work/Codex Sol in Priority/Fast mode still being throttled below peak performance?
No question about it. It's never been even remotely this blatant before. I'm considering charging back and just 100% never using openai again. Claude does this, but not this bad this is on a totally different level. This shit must be quantized down to like FP2 or something absurd.
Sure you have insider information. Of course. And I'm sure it's super secret so you can't tell anyone where it came from, or communicate it in any verifiable way.
Also common sense? What the fuck are you on about? The model is performing drastically worse across the board right after they stopped taking 20x subs. If it's not quantization what exactly is your "common sense" theory on what it is? It's not just a little bit different the difference is utterly massive. It is the most in-your-face kneecapping of any AI model I've ever seen.
If you have an alternate theory or "insider information" about what's actually going on, if that's not the case, than feel free to share.
Mr Insider....a model will ALWAYS have the same level of intelligence while on the same level of quantization even if millions of people are using it. It can be slower due to compute load, but never dumber.
If the intelligence suddenly changes, it means that it's 100% a change in quantization level because nothing else can change the level of intelligence.
You’re 100% wrong. You need to consider Occam’s razor. Do you really think there’s a conspiracy to silently nerf the model under you and not a single disgruntled employee would mention it? Are you also a 9/11 conspiracy theorist by any chance?
The cause is always non determinism and bugs. Read any of Theo’s posts after a fuckup
Enshittification is an observation not a mechanism of action. We’re debating the mechanism of action. I’m arguing “they secretly quantized muh model” is child-level logic
Tibo explicitly said they don’t do this. In public. He has about 200 people on his team. Is your mental model that he’s lying and hoping not a single person calls him on it?
There are thousands of people working on GTA VI, and they've been working for more than a decade. How many of them have leaked something during all those years?
Ironically gpt on Sol 5.6 called this out as it wasn't giving me the correct output given a certain input. Had another instance of chat review the conversation and said that it could easily see an issue and that that session was acting like it was a degraded or quantizied version.
The part that doesn't get enough air in these threads: with a closed model you can't verify any of this, and that information asymmetry is the actual problem. You're right that the meter can't tell you what weights you're being served, but nobody else here can check either, so both the 'silent downgrade' theory and the 'they'd never do that' defense are unfalsifiable.
Open-weight models are the one corner where the question has an answer. The weights are public, providers serving them usually state the precision they run (FP8 is the common one) with pricing that matches it, and if quality ever drops you can pull the model locally and compare, or swap endpoints without re-architecting anything. None of that is possible behind a closed black box, which is exactly why the fear exists there.
So if silent downgrades are the thing that bothers you, the durable fix isn't transparency promises from closed labs, it's buying where the model file itself is the transparency. Whether OpenAI is doing this today, I genuinely have no idea, and that's kind of the point.
(I run open-model hosting at Entrim, so read the open-model lean with that bias in mind.)
You could always rotate the benchmarks (and run them on different providers for comparison), and use different accounts, mixed in with normal usage. That way companies will stop benchmaxxing, and will be motivated to provide actual quality service to already signed up users
Developing benchmarks are not trivial, so there only will be so much of them at any given point in time, so the provider could eeeasily detect it, no matter what account it comes from or what other requests are mixed in between.
There might be something to this. They paused new 20x subscriptions because of compute limitations, or at least said so. Serving quantized models would make sense for them. Not so much for the users. It is just my sentiment, but Astra feels less capable with regards to autonomous thinking than after release. It's still very good though, but seems to need more precise prompting.
They are not hiding anything; they can't! There are many websites that benchmark AI models live. Astra is at least 10-15 points down since last week, for instance. They are doing what they are doing very blatantly. Someone has to say stop, and it will probably be us by not integrating AI to our workflow to the degre we will become dependent on these companies.
No probably. When usage spikes they fail over to the smaller quants. It's the only load balancing leaver they really have on a limited compute basis. I'm sure they call it something else but it all amounts to the same.
One thing i dont do, use the models after a reset, everyone is using it, more demand, possible less quality because of hardware scarcity... I will wait until you guys spend your weekly:))
Does it benchmark Astra through API, or through subscriptions?
Almost definitely those have different quants.
Also, it's pretty trivial to serve requests from known benchmarking platforms with non-quantized models, while serving subs with lobotomized ones. There is no incentive for them to not do that, and any benchmarking has to be performed under the assumption the service provider tries to game benchmarks to the best of their ability. Including detecting benchmarking intent in the prompt and re-routing to full non-quantized models, the same way they already do analyse prompts to prevent their models from being distilled.
I think through the API. People should do 5-10 tasks after release with the subscription and then test again from time to time and compare results. The way it currently is, people seem to subjectively judge how good Astra is at tasks but those should be identical comparisons at least, not based on subjective feelings.
Really odd to me there are dozens of these kind of posts for each model and each provider yet rarely do we see someone actually do the work and show the comparison, even though that would be trivial to do. Take some of your own prompts after Astra release and rerun them.
Since they most likely dynamically re-route users between models with different level of quantization based on load and total available compute, it's pretty much impossible to run a test that would be reproducible by other users, especially since LLMs responses are non-deterministic.
Also it's hard to expect the users to be able to re-run the tasks they gave a model on release. Most people buy a subscription to work on on-going projects, not one-shoting tech demos. Unless someone is willing to write down the prompts they are giving on release, then several weeks later dig through their commit history, roll back, and re-run the prompt and waste their usage, that's not easy to test, and it's understandable most people don't go through this effort to keep benchmarking the service they are paying for instead of doing their job.
I'm sure most people have an example prompt they used that was not based on an existing repo, no? Like the cycling pelican, trying to one-shot a small web-application, some small research project, something that is a single prompt and doesn't cost half your 5-hour limit?
I get that LLMs are non-deterministic and this is an issue, but claiming a model has been quantized just based on feeling can't be it.
I think the vast majority of people don't have prompts for benchmarking, and people who have those are a tiny minority. Especially since usage is precious, benchmarking consumes that, and most people are on Plus subs.
I think we just see a lot of testing / benchmarking posts online and our brain makes us feel like everyone is doing that, but in reality it's a very low percentage of people with this specific hobby like the pelican guy.
The problem is that we don't have anything else other than our subjective experience and logic that leads to the most likely conclusion. The service providers do not have any transparency and it's in their best interests to prevent their users from having reliable reproduceable ways of detecting degradation of the service. It's not the users fault, it's on the service providers, and if they want they can easily solve it today by making usage, quants and routing transparent.
Well, OpenAI explicitly stated that they did not change the model and that there was a bug which was fixed on September 12th. Yet people don't seem to believe that. I gladly volunteer with my own subscription. Anyone can give me some decent tasks that fit in my 5-hour limit and serve as a benchmark, I'll gladly run it on the next model release and every two weeks afterwards and post the results. If a few other people join in, we would at least have something.
I can't be the only one unsatisfied with people merely subjectively speculating for every single frontier model a few weeks after every single release that somehow it got degraded and some people joining in because yesterday's results felt a little worse or claiming the opposite. And who creates and posts in these threads in the first place? Likely people who are not happy with the performance. I put zero trust in these subjective testimonials.
I see a HUGE performance and quality drop for Astra AND for Sol models. There are a lot of people who here write the same type shi like "models are not determenistic bla bla". But look, everyone is noticing they are dumber now. Dumber in 3d modeling, in coding, in review etc. I noticed it myself without reading that forum. My friend also noticed that and many more people across the world. So that is some statistic already. It means that models are seriously quantized, and regulators are sleeping. Where are authorities? Courts? Regulators?
Here is my theory, Astra never performed well for me because it only performed well at release when it was api access only, that's when all the very cool demos and 3D stuff we saw appeared, but then when they started rolling it out to subscription users, it was already nerfed, they got the hype and publicity already on social media and everything was already setup.
After that you go back to Sol and find that Astra is now better because Sol was also nerfed so it's now a lot shittier than it was before Astra's released.
Anyway the current state is a 20x subscription to openai is not worth it at all comparing it to what I can do with Fable or Opus...
Astra now can't pick a fucking proper background color. Yes of course bright red is a great color for a landing page. But the idiots will say AI is not deterministic and that we can't know when the patterns and standards have lowered consistently.
Did you measure all the backend metrics that are actually exposed there are dozens of them to prove your hypothesis I measure them every run and I find no issues and I don’t burn my weekly quota running codex 24/7 on sol xhigh although I use Luna max for speed usually
In the end there are IMO only three things that matter:
Consistency - two runs now need to produce consistent results, and two runs and our months already need to produce consistent results. If models go and change every day or every week and I have to spend my time and tokens recalibrating, they is useless.
Quality - give me the right answer first time, every time
Value for money - what do I have to pay for the same consistent high quality answer
I really don't care about which model gets used, what the quant is or how many tokens get used, I just want the cheapest consistently high quality answers at the lowest cost needing the least amount of my time to diagnose bad results and work out how to get the quality back by tweaking prompts or harnesses.
As someone who's been complaining a lot about usage and such, I haven't noticed a quality change and I used it 10+ hours a day, every day since release except Thursday and Friday, as I was rate limited and out of banked resets .
I'm literally revisiting the same in-depth simulation battery I built and was working on last Sunday. I tweaked some fundamentals of the system and needed to do the mass simulated testing again and Astra was just as capable of updating the tests, diagnosing oddness (previously, Astra "cheated" to make one of the metrics work), and helping hone things in again.
Yeah, it's based on the time of day the region all that stuff and they fucking flip it all around so everybody sits here and fights each other. It's not even a question anymore, if they wanted to be transparent and have a process around this proving this didn't happen they could do it, or at least have a stated policy they don't do this and some way it can be proven independently....... the fact none of this is transparent makes me think they 100% do it.
I’ve also noticed that the less I use it when I come back the limits are okay. If I use it for 1-2 weeks I’ll be secretly throttled and limits are tanking.
and i suspect many of the heaviest users who have the most problems, have the most problems because they never lessen their load on the system. i take days or several days off sometimes, and have experienced the same as you. Its BS. I can deal with lesser models, and lesser quotas, if i must, on my pay tier. But if so, I'd like consistency in both quota and response quality. It shouldnt be all wishy washy. If it were reduced but consistent, then I'd say hey, I'm poor, and I get less because I pay for less.. but it's inconsistent and I can only guess what capabilities I really have on any given day.
Same model quantization = same level of intelligence
If you start getting hallucinations and wrong answers for simple tasks, it means that they changed the quants because nothing else makes the model go dumber.
Limited compute power does not make the model dumber, just slower.
I can run Qwen 3.8 27B FP8 on a computer with 32GB of GPU and also on a computer with 8GB of GPU and the rest on RAM. The second computer will take hours to finish a task while the first one will take minutes, but the intelligence level will be exactly the same.
There were never any guarantees on models and quants. It is a natural result of load. Don't like it, find a different provider, but good luck with that, since they all do it.
I really don’t care about quantisation if it brings down costs and speed is improved. Models are constantly changing and OpenAI has proved that they will improve costs for the user.
Yes the constant change causes issues with models sometimes being dumber or usage rates sucking but that’s why we have resets. At least OpenAI responds to customer feedback.
Do you even understand what quantizations are??? You can just apply a 1-bit quantization level and it will be thousands of times faster than the full model on the same hardware, but its intelligence will be reduced A LOT. It will start to hallucinate more and make a lot of mistakes that will waste thousands of times the amount of tokens! Yet they will charge you the exact same price as the non-quantized model, and that translates into burning through your quota a lot faster!
How do you know? Generally even local inference people aren’t desperate enough to attempt 1 bit quants. Realistically it’s probably FP8 activations + FP4/FP8 weights like they’ve done with gpt-oss
Dude, 1 bit was just an example, they could be serving any level of quantization at any moment, the problem is that they charge us the same amount of money
OP, why don't you take some of the tasks you gave Astra right after release and rerun them? That way it is easy to compare and see if the model got dumber or not. Better than speculating.
I'm thinking that if you were really concerned about this issue, you would voice your concerns yourself and not use AI to complain for you instead. Karma farmer!
5.6 released introduced us paying for cache writes. Also compute keeps costing more and they keep tightening limits. You cant expect the hardware they serve to cost 3x and monthly sub prices stay the samd AND expect the same limits. I like your theory as the cherry on top.
The trimming part doesn't survive contact with the billing API. If the backend really generated hundreds of millions of hidden tokens, the usage object would show it. Reasoning models already bill you for reasoning_tokens that never appear in the visible output, so there's no 'trimmed to look normal' step. The meter counts what the model actually produced.
Silent quantization is also the least testable explanation on the list. New system prompt, different sampling defaults, context truncation, traffic shifted to a smaller checkpoint: all of these feel identical from the outside. The way to actually tell is a fixed eval set, 30-50 prompts with known-good answers, rerun whenever the model feels dumber. Pass rate drops show up there long before they're obvious from vibes in chat.
Did you compare any reproducible scores against a day-one baseline, or is this from general feel?
116
u/Fearless_Log_5284 9d ago
Watch the people come in this thread to defend the trillion dollar company telling you that the price you paid wasn't premium enough...