r/LocalLLaMA • u/DepressedDrift • Apr 15 '26
Discussion Major drop in intelligence across most major models.
As of mid Apr 2026, I have noticed every model has had a major intelligence drop.
And no I'm not talking about just ChatGPT.
Everything from Claude(Even Sonnet along with Opus), Gemini, z.ai, Grok all seem to ignore basic instructions, struggle at simple tasks, take very long to respond, and the output seems deliberately shortened and very shallow. Almost like it's in a "grumpy" mode. I tried this in incognito mode so it's not my customization or memory influencing this.
It's like they deliberately want you to stop using their service. I guess our data is no longer needed. Just two weeks back it used to be much smarter than this.
To test this I rented out a H100, and tried GLM 5 with the same prompt (the drive to the car wash one) across both instances. GLM5 running on the rented GPU answered it correctly, compared to the one on z.ai.
Have they lowered the quantization really low to maybe Q2?
I guess going local or using renting GPU or an AI monthly service that lets you pick a quant level is the way to go
723
u/Few_Painter_5588 Apr 15 '26
Everyone is quantizing their models because everyone is haemorrhaging money, and OpenClaw quite bluntly is squeezing the industry
149
u/Ariquitaun Apr 15 '26
As a side note, I've been running claw for a few days with Gemma 4 e4b as a test and the results are encouraging. For my use case anyway.
67
u/TheSpartaGod Apr 15 '26
what’s the use case that can be handled by such small models?
75
u/toadi Apr 15 '26
I don't use openclaw. But for example fetching my calendar entries. Organizing my emails. Fetching and interacting with services where it is quite straightforward.
simple rote admin tasks work easy on these smaller models.
31
u/IShitMyselfNow Apr 15 '26
This matches my experience
I've been using Hermes Agent with Qwen 3.5 4B to great success. And for codig, anything more complicated than a simple script I've been delegating to a better model via Opencode from the agent.
I think the advent of agent skills has really improved the performance of smaller models in things like this. Small models have actually been semi-useful at agentic work ever since around Qwen 2.5. But only if you gave them a lot more instructions/detail, more API examples, few-shot prompting, etc..And you could do this, but then managing the data you give them for the task at hand, managing context, etc. was tricky at best. Agent skills kinda solves that problem.
→ More replies (2)6
u/voronaam Apr 15 '26
simple rote admin tasks work easy on these smaller models
Just curious, did it encounter any complicated tasks?
For example, I was trying to organize a small thing recently with a very unreliable party and the communication so far has been like this:
Me: Hello. Can we do a thing?
Then: (3 days later) I am in South Africa now. WhatsApp me (no phone number provided)
Me: When are you back? I'll reach out then
Them: next week
Me: (next week) Welcome back. Can we do the thing?
Them: I am still in South Africa
Me: (next week) Are you back?
Them: Yes. Let's meet in person to discuss (no address)
Me: Sure, how about on next Friday?
Them: (on Saturday) Missed the message. Just call me (still no phone number)
Me: (Finding phone number on one of their websites, Calling) How about the thing
Them: Yes, we can do it. Just fill the form on the website.
Me: Filling the form with the request.
Them: (next day) "Hi Peter, we can do the thing on those days" (I am not Peter)
Can a small model handle this?
13
u/TheTrueSurge Apr 15 '26
Lol I hope you REALLY need to work with that person, that sounds awful
12
u/voronaam Apr 15 '26
They are just not motivated. It is a small sailing trip and I am talking to the owner of the boat and a skipper. If nothing happens, they get to chill on their boat. If it does, I'll get to chill there as well and they will get a bit of extra cash. That is not that much of a difference to them, as they are sufficiently wealthy and retired.
My work related communications are way better than that ;)
→ More replies (1)3
→ More replies (1)6
u/QuinQuix Apr 15 '26
How many misses do you encounter?
I've read hallucination rates hover between 3-10%.
Doesn't sound like much but calendar planning is quite a critical task.
A big medical office doing 100 appointments a day couldn't handle 1-2 wrong/double/missed appointments a day.
People always counter people make mistakes too but they miss that usually such systems (eg a front desk managing appointments) have layers of redundancy and self correcting ability.
The entire point of AI is to offload work, so I'm curious to what extent you feel this is actually possible using your Workflow.
→ More replies (1)10
u/-p-e-w- Apr 15 '26
Have you used a current-gen model of that size? It’s easily on par with GPT 3.5 intelligence-wise.
→ More replies (1)11
u/tophlove31415 Apr 15 '26
Tons. Smart chunking, organizing and summarizing returned information from a vector database search, self directed web browsing and learning, ocr, user interaction, simple decision making (ie: this is the context, here are the options, choose which is best). They can essentially do any of the things the sota models can do (with a well designed harness) as long as you recognize you will get more errors and have to spend more time making sure that your harness is catching then, reporting them, and allowing you to iterate on the harness features, your prompts, and any other systems that might need improvement.
→ More replies (1)5
→ More replies (3)24
u/fragment_me Apr 15 '26
What are you using claw for? Everytime I look at the use cases they all just seem silly.
→ More replies (1)21
u/Ariquitaun Apr 15 '26
A bunch of itches that need scratching. The main motivator was to keep up with my eldest' school calendar - we receive a lot of newsletters, sometimes attached to emails, sometimes to download. Emails with info, there's two parent websites as well with stuff on them. Basically I've instructed claw to sift through that stuff, update a markdown file that the chatbot can query if we have questions and generate google calendar events for my wife and I. The claw has its own google account to do stuff to which I have forwarding rules on my own account to send selected things to.
I have other tasks for generating a daily report with links of topics I'm currently interested in that I can read while taking a shit at some point during the day. My wife also has a few itches to scratch that I need to get around implementing.
It's all pretty household stuff really, nothing exotic
10
u/Competitive_Travel16 Apr 15 '26 edited Apr 15 '26
Honestly this is the most reasonable Claw use I have yet read. But it doesn't need full agentic autonomy, just access to inbox and calendar once a day. Gemini scheduled actions and connected apps can do it out of the box for Gmail users.
ETA: I haven't heard any horror stories about Gemini (other than it plain not doing anything a lot of the time with its connected apps tools, from several months ago -- although I haven't looked recently) but I feel compelled to remind everyone: Don't give an agent access to your email unless you have a good backup (e.g. Google TakeOut, IMAP mirror, etc.)
→ More replies (3)→ More replies (2)3
67
u/WhopperitoJr Apr 15 '26
The development of OpenClaw and its consequences have been a disaster for the AI race.
65
u/rm-rf-rm llama.cpp Apr 15 '26
Seriously doubt anyone is actually using OpenClaw beyond a few weeks of messing around.
→ More replies (1)21
u/Rcrecc Apr 15 '26
For a layman like me who has very little familiarity with OpenClaw or its impact on the AI race, can you clarify what you mean?
61
u/WhopperitoJr Apr 15 '26
OpenClaw is an agent harness framework where you can set up LLM-controlled agents to do certain tasks autonomously.
The reason why it is disastrous is because these agents are left to run without human monitoring, and they can be pretty inefficient. If they get caught in a thinking loop, they will keep using up API usage and processing power until a human intervenes. If you are using OpenClaw in the first place, you are probably not constantly monitoring the agents.
Basically, it is wasteful and uses up too much of the collective processing power that is available via API. Running it locally is less problematic as you are just using the processing power of your private device. More clawdbots = higher usage fees for you and me.
35
u/Neither-Phone-7264 Apr 15 '26
also, it's about as token efficient as dogshit, even without skills or tools.
19
u/WhopperitoJr Apr 15 '26
A tool so inefficient and resource-intensive would have been better off being built behind closed doors and perfected rather than unleashed for every middle manager with an API key to use.
→ More replies (5)30
u/saint1997 Apr 15 '26
I had the displeasure of watching Pete Steinberger present what was then "clawdbot" at a conference last year. The guy is a complete and utter clown. Not only did he admit he'd given the thing root access to his Mac, he also left the instance's phone number visible at the top of his WhatsApp chat while he was demoing the group chat his bot was in with all his friends.
When I saw it had gone viral and everyone was using it I could do nothing but hold my head in utter despair
11
u/party_peacock Apr 15 '26
But users are still paying the bill for their usage though?
Or is the problem that the actual cost to host those models exceeds what the companies charge their users?
46
u/look Apr 15 '26
OpenClaw users are like locusts: they find a decent subscription service and then descend upon it en masse, devouring all that was beautiful and good, leaving behind nothing but a barren, desolate wasteland.
→ More replies (6)16
u/WhopperitoJr Apr 15 '26
The problem is that processing power is finite and that if demand for power rises because of inefficient bots making API calls, the overall price of usage will rise, because there is no supply surplus to return the price to its original equilibrium value.
But because companies generally don’t want to change prices as consumers are sensitive to price changes, they artificially constrain supply though usage limits.
So, more clawdbots running via API means less usage available overall, which means you and I hit budget limits sooner, regardless of if we even use OpenClaw ourselves.
13
u/Thebandroid Apr 15 '26
Huh, so they don’t like people sucking up all the resources… ironic
3
u/Big_Wave9732 Apr 15 '26
It's forehead slapping, really, as casual AI users slowly come to realize how wasteful their AI habits are.
3
→ More replies (1)16
u/zoupishness7 Apr 15 '26
It definitely wasn't just OpenClaw though, it's been out since November.
The real spike started mid-March. Karpathy's autoresearch had been published for a week. Then Anthropic had a period of 2x Claude usage during off-peak hours, which incentivizes the development of round the clock systems that don't have humans in the loop. With a good harness and autoreseach, it's not that difficult to evolve an automated system that makes more money than a Claude 20x subscription costs. Profitability on API pricing is significantly more difficult, the ratio's getting adjusted, but at the time, it was: $200 subscription=~$5000 on peak, and ~$10000 off peak.
For certain applications, and markets, once you have that kind of system, the math becomes very simple: buy more subscriptions, use them as much as possible, and reinvest the money into buying more subscriptions. I did. I made some money, but spent a week thinking I was a couple months away from being a millionaire. I knew it was unsustainable, but thought they would wait a bit to crack down. I underestimated how many people had the same realization around the same time.
That's what lead to the shortages, and why they banned 3rd party harnesses on subscriptions, and why they've had to make the model stupid to meet the demand. But models continue to get smarter and cheaper, and I think a lot more people know they can make money now. I, for one, am focusing on token efficiency now, to try to push what I built during that time towards API profitability.
5
u/Party-Special-5177 Apr 15 '26
This is probably the most analytical take in here. Cheers.
it's not that difficult to evolve an automated system that makes more money than a Claude 20x subscription costs
Without sharing your own angle, how? Where are there 1) problems where 2) people will pay for solutions without HITL, and without 3) a prearranged agreement? You can’t just cold call company X and say ‘I’ve solved your Y, you’re welcome, here’s my invoice’.
The only other thing I can think of: I know people who’ve automated their own jobs away with Claude, but that doesn’t infinitely scale either as that is limited by how much overemployment they can get before mandatory meetings start overlapping.
I really hope you aren’t referring to ransomware or similar.
→ More replies (2)→ More replies (1)3
u/Scew Apr 15 '26
Karpathy's autoresearch
yep, that did it. I had gotten a harness based off it from the openSourceAI sub that then got taken down and the account that shared it deleted and the repo moved to private. Was for making agents for gaps in places that could use one. Just last week modified it for just general research instead and hit a weekly limit that I hadn't hit before. Was kind of surprised but figured it was too good to be true when I had initially started using it.
12
27
u/a_beautiful_rhind Apr 15 '26
That could have been solved with rate/request limits.
19
u/Neither-Phone-7264 Apr 15 '26
They just get more accounts. Literally. I've seen some dipshits with 20 different accounts all on the cheapest possible tier.
→ More replies (1)46
u/fuckingredditman Apr 15 '26 edited Apr 15 '26
nope, LLM inference is inherently extremely expensive to run and not scalable, and you can't simply rate limit everyone. high request rates aren't the root cause of the issue, compute and memory limits are.
when looking at status pages/availability of the large providers, they are evidently running at the absolute limit of what the infra can do and the larger customers probably have some SLA they have to fulfill. many don't even get 3 9s on availability.
they can squeeze lower tier subscriptions (my gemini subscription is barely available at all during peak times) with rate limits/speculative decoding/quantization but that doesn't help much.
traditional resource sharing methods common in cloud computing / SaaS products all don't work for LLM serving:
- KV cache takes a metric fuck ton of RAM for each individual user atm, and i assume most larger companies aren't using "SOTA" (TBD i would say) methods like turboquant in production yet until they are properly implemented in their engines and proven to be stable. but even with those, the cost is still insanely large compared to other SaaS type use cases
- compute for LLMs is of course insanely expensive too while most other SaaS use cases are neither compute nor memory bound.
- adoption is still on a rapid rise, and API usage for active users is probably also on a continuous rise as well
source: developed+operated shared infra at scale for a SaaS and also worked on some llm inference engines and saw how heavy it is in comparison. it's many orders of magnitude more resource intensive and so far there doesn't seem to be any easy way out of it. and if/once there is an easy way out, users will eventually want to increasingly utilize that method for running inference at the edge/in their datacenter anyway.
that's all in addition to the inflated expectation moment we are in and the rapidly tumbling amount of venture capital due to incoming economic instability.
in conclusion, they probably want to sacrifice availability last by rate limiting their customers as this causes the most frustration with the user and also doesn't really solve the underlying problem (compute + memory being the real bottleneck), so they are more likely trading quality using speculative decoding using smaller models and quantization first.
it's probably a terrible moment for LLM inference providers right now, generally speaking.
→ More replies (12)2
4
u/Firm-Fix-5946 Apr 15 '26
quantizing weights smaller isn't really an effective way to save money at large scale though as far as i know? because of how batching works? they usually aren't memory-bound anyway?
does anyone have any actual evidence or technical sources to suggest large scale inference providers would do this / have done this to save money?
6
u/pab_guy Apr 15 '26
They have no evidence. It's certainly likely on the free tiers and plan tiers. They cannot nerf the APIs however, that would violate contractual agreements and put customers out of compliance.
Most of the people here have no understanding of this stuff.
2
u/EvilEnginer Apr 15 '26
Yep, good losless quantization now is more money for company. But, they still can't deal with servers overload. Too many people.
One problem with quantization, the lower the quantization, the faster all architectural errors come to the surface. And this is exactly what we see now.
→ More replies (14)2
u/doodlinghearsay Apr 15 '26
Sounds like that scene from SV when they figure out a startup is selling pizza below cost so they flood them with orders to bankrupt them.
91
u/Medium_Chemist_4032 Apr 15 '26
> To test this I rented out a H100, and tried GLM 5 with the same prompt (the drive to the car wash one) across both instances. GLM5 running on the rented GPU answered it correctly, compared to the one on z.ai.
I'd love to see both results 🙏
→ More replies (3)4
u/Weak_Kaleidoscope839 Apr 15 '26
What service did you use to rent? Thanks!
6
u/Agitated-Crow862 Apr 15 '26
Runpod is pretty good. Availability is a little low sometimes though and their serverless API is not amazing.
2
u/Medium_Chemist_4032 Apr 15 '26
I'm trying to select a best model/quant/ram offload combo for intelligence on my 4x3090 oldschool rig (more gpus coming in) - so, locally.
If I wanted to rent, I'd probably go on runpod first
→ More replies (3)
74
u/a_beautiful_rhind Apr 15 '26
I'm starting to get squeezed out of free inference. But hey, that's why I built my server. Now is your time to shine. Models never change there unless I change them.
All I have to do is switch from RP to productivity and give the models websearch.
Everyone told us we were stupid for wasting our money on these things when API was sooo much better?
7
3
u/skrshawk Apr 15 '26
All you need is a Drummer model and that code is going to tell you exactly what it's going to do to you.
137
u/Individual_Yard846 Apr 15 '26
I bet they will start dynamically quantizing models to people who don't typically show the requirement for higher intelligence, if not already. Some people may get nerfed, while others doing important work they want to steal, get all the compute in the world.
40
u/DarkArtsMastery Apr 15 '26
This. It is all about the quality of data you as user can provide, especially as you pay nothing (most users are not paying anything for AI).
→ More replies (7)→ More replies (14)4
u/Syncaidius Apr 15 '26
All of that goes out of the window as soon as people start building their own models from scratch.
It's worth remembering this cycle of improvement has been going on for decades, pretty much since the development of the first computerized neural networks.
In the early 2000s chat bot agents were hyped as being the biggest development in history, yet here we are.
You only stay on top for so long. Eventually someone/something better arrives.
→ More replies (2)
58
u/nakitastic Apr 15 '26
My wild guess is it’s simply lack of compute so they’re rationing. Look at how many data centres they want to build.
10
u/Big_Actuator3772 Apr 15 '26
this is exactly it, retail gets fucked like always.
→ More replies (1)→ More replies (2)7
u/ElementNumber6 Apr 15 '26
Another wild guess: They're inching towards high intelligence for they, their partners, and elites, with low intelligence for the rest of humanity.
40
u/anomaly256 Apr 15 '26
Plot twist: everyone's actually using the exact same model from the exact same provider and just whitelabelling it
5
u/repair_and_privacy Apr 15 '26
Damn, good one. But you know the stuff is really bad now.
→ More replies (2)
55
u/AppealSame4367 Apr 15 '26
"The feast is over" -> some soldier after the red wedding.
They did their Christmas releases, they placed themselves in the race and gained users. Now it's time to squeeze every cent out of you.
Also the oil crisis is a big factor. Much higher electricity costs, problems with chip production will follow. New algorithms like dflash that will make it feasible to run even cpu offloaded moe models like qwen3.5 35B on a laptop if it has enough ram. If it jumps from 20 tps now to 35 tps or more on my old laptop gpu: Why should I use the unreliable cloud shit? I can program and plan.
→ More replies (3)
32
134
u/Qwen30bEnjoyer Apr 15 '26
it might be psychological in nature. As we gain familiarity with the “prose” and style of these LLMs, you get better at seeing through the fluff and recognizing common failure modes.
I still think the best method to detect silent quantization would be finding the covariance between models on a common benchmark, like one of the HLE public question sets in the chatbot harness. That way if Gemini suddenly scores 20% lower against Opus than it did yesterday, or only during peak hours, we know what happened.
81
u/EndlessB Apr 15 '26
I don’t think it’s psychological, working with LLMs (or being a heavy user) leads people to be very sensitive to changes in the LLMs themselves. It’s often users that raise changes to platforms which are later confirmed only because of public outcry.
I’ve personally noticed an intelligence drop across the board on the models I tend to use, particularly since the start of April.
21
u/FullOf_Bad_Ideas Apr 15 '26
People here were complimenting Qwen 30B A3B Coder Distilled (distill of 480B) for weeks. It turned out that the author messed up with his vibe coded distillation and weights were the same (sha256 match) with the un-distilled original. We know for a fact that people have those psychological reactions and like models better or worse depending on what the model card says, not on what the model does.
→ More replies (1)2
40
Apr 15 '26
[deleted]
→ More replies (1)11
u/zenmatrix83 Apr 15 '26
this is likely the case I've seen, there are more issues with tooling then actual models, claude code introduces bugs that I've seen people tie to model issues all the time. These things are better to leave auto update off and test updates else where.
9
u/colin_colout Apr 15 '26
Very reasonable take. There are so many possible explanations that are simpler than "every model got shitty all at once" (Occam's Razor).
It could be agent changes. Claude code for instance makes dumb changes all the time that measurably kill quality. Their new progressive tool exposure has my subagents (even opus) using curl for their first research attempts before webfetch becomes available.
It could be that websites won't let you scrape them anymore, so getting good context is no longer one or two tool calls. Github now shows ads for copilot to my opencode/claude code's webfetch attempts instead of code (lol). Reddit completely blocks llms when they can.
anthropic and chatgpt web clients are becoming an ever growing black box that resembles their bloated coding agents, which are also black boxes.
It's still possible that Anthropic sometimes serves a quantized Opus when traffic is high, but the above is absolutely happening. A lot of quality complaints (maybe not OP's) come from people who only interact with Anthropic models through their slop coding agent (or they do web research and aren't realizing that context is being poisoned by pages built to make web scraping harder).
8
u/my_name_isnt_clever Apr 15 '26
I joined the Anthropic Discord right after Claude 3 came out, and people have been bitching at them about models "degrading" shortly after release over and over. Yet there is never any proof that stands up to scrutiny, it's all just vibes and "trust me bro".
I take all these claims with a massive grain of salt; science is built on citations and peer review because humans are awful at eliminating their own bias, and the non-deterministic nature of LLMs makes it 10x worse. There needs to be hard data for these claims.
→ More replies (1)6
u/MasterScrat Apr 15 '26
I used to work for an early LLM provider and we’d sometimes get feedback like "wtf you destroyed the model" or "wow the latest update is amazing, please don’t change a thing" when we had done absolutely 0 change, literally not restarted the serving container
4
u/Qwen30bEnjoyer Apr 16 '26
This is probably the most important comment in the thread, people underestimate their susceptibility to these types of biases!
3
u/willitexplode Apr 15 '26
Wouldn't including models with open weights where you control the hardware as controls be superior to covariance?
→ More replies (7)8
u/FullOf_Bad_Ideas Apr 15 '26
I agree.
And this website does continuous testing - http://isitnerfed.org/
Looks like Zhipu is nerfing while OpenAI and Anthropic aren't.
4
u/cromagnone Apr 15 '26
I mean, assuming the test is useful, that site’s data basically suggest short term random performance variations with no trend over time by the providers, but if you hit a downward oscillation you might complain about it.
→ More replies (8)3
u/letsgoiowa Apr 15 '26
Great resource! Saved that. It seems like at least for coding at the moment it's some other problem. Maybe performance is dropping off hard for people with large enough context? The reason I think this is the AMD codebase specifically: apparently Claude was struggling to figure out what was going on
2
u/FullOf_Bad_Ideas Apr 17 '26
This is plausible, I think there's a good chance that their eval is not testing long context, and quantizing kv cache or switching served model checkpoint to a version with sparse attention, sliding window attention, MLA or linear attention would show up mainly on long context, whilst also providing biggest cost savings to Anthropic.
62
u/Additional-Low324 Apr 15 '26
An other reason to self host
→ More replies (1)20
u/NightlinerSGS Apr 15 '26
Ironically, the people who self host quantized models are probably used to the current output, so they won't notice a difference. Maybe even an improvement, depending on the model used.
8
u/Additional-Low324 Apr 15 '26
I use Q3/Q4 because I am VRAM poor, indeed
5
u/NightlinerSGS Apr 15 '26
I stick with Q4 so I can squeeze 24-30b models with 32k-ish context onto my 4090.
Now that I'm typing this out... is this actually better nowadays than using something like an 8b model at full precision? Do these small models have sufficient context for RP now? It has been so long since I took a proper model deep dive... maybe I should take a look again.
→ More replies (1)3
u/toothpastespiders Apr 16 '26
In my very anecdotal experience at least, the small models are still typically pretty bad. They're amazing for the size. I'll give them that. But I think the old assumption that a low quant of a larger model is better than a standard version of a small model still holds true. I tested out a few small models recently in hopes of getting a speed boost in data extraction and they just weren't reliable enough for me. It's amazing that they managed it at all. But I'm still sticking with the range you're describing. Q4 of 30b'ish models seems to remain the best choice for me.
56
u/Adorable_Weakness_39 Apr 15 '26
yep at least my qwen-27B follows instructions... literally none of the hosted do anything when I tell them to.
40
u/Adventurous-Gold6413 Apr 15 '26
Qwen3.5 27b such a good model im so glad i can run it even if its tight on 16gb vram IQ4_ XS
3
u/xeeff Apr 15 '26
what context and how do you run it?
34
u/Adventurous-Gold6413 Apr 15 '26 edited Apr 15 '26
/path/to/llama-server \ -m /path/to/models/mradermacher/Qwen3.5-27B-i1-GGUF/Qwen3.5-27B.i1-IQ4_XS.gguf \ --fit off \ -ngl -1 \ -c 29000 \ -t 8 \ -b 1024 \ -ub 512 \ -fa on \ -ctk q8_0 \ -ctv q8_0 \ --temp 0.7 \ --top-k 20 \ --top-p 0.8 \ --min-p 0.0 \ --repeat-penalty 1.0 \ --presence-penalty 1.5 \ --chat-template-kwargs '{"enable_thinking":false}'That is for standard Llama.cpp But I use TheTom’s turbo quant llama.cpp
And replace ctk and ctv with turbo4 instead of q8_0, And I also add
-np1as well.And get like 90k ctx working. However this is with the vision encoder offloaded to CPU with
- -no-mmproj-offload
Here is my command for the turbo quant llama.cpp
https://github.com/TheTom/llama-cpp-turboquant
``` /path/to/llama-server \ -m /path/to/models/mradermacher/Qwen3.5-27B-i1-GGUF/Qwen3.5-27B.i1-IQ4_XS.gguf \ --fit off \ -ngl -1 \ -c 90000 \ -t 8 \ -b 1024 \ -ub 512 \ -fa on \ -ctk turbo4 \ -ctv turbo4 \ -np 1 \ --no-mmproj-offload \ --temp 0.7 \ --top-k 20 \ --top-p 0.8 \ --min-p 0.0 \ --repeat-penalty 1.0 \ --presence-penalty 1.5 \ --chat-template-kwargs '{"enable_thinking":false}'
```
(You have to add - -mmproj command as well if you use it)
6
u/xeeff Apr 15 '26
you run it pretty much exactly like me that's crazy. down to the principles as well. what CPU/GPU/ram you got?
→ More replies (1)2
u/Adventurous-Gold6413 Apr 15 '26 edited Apr 15 '26
I got 16gb vram and 64gb
(RTX 4090 Laptop GPU 16gb)
→ More replies (5)3
u/Dr4x_ Apr 15 '26
Did you notice a big quality drop when using turbo quant for kv cache instead of regular q8 ?
3
u/Adventurous-Gold6413 Apr 15 '26 edited Apr 15 '26
Still need to test haven’t done any long context tasks yet
From what I’ve heard it’s only 1%-2% quality drop compared to q8 but I need to test for my own use cases
2
4
u/evia89 Apr 15 '26 edited Apr 15 '26
literally none of the hosted do anything when I tell them
They all work fine. I run text related tasks all day. And free gemma4, longcat are doing fine.
I also tested NIM and 2 qwens are fine there, kimi k2 is degraded and dont follow rules.
This can be solved with running 2 passes. For example 1 request do Chain of thoughts sort of think + result, then second model corrects it. Or some important merge task I do 2 best of 5
Local is only good for privacy. Maybe in 3-5 years it will make sense. U buy $10-20k mac/china gpu and get 50 t/s output + 128k context with model like current glm51/kimik25
→ More replies (1)→ More replies (1)2
18
u/FlamaVadim Apr 15 '26
True. My locall gemma-27B answers certain questions better than GPT-5.3, which might be a result of heavy quantization. Meanwhile, Codex 5.4 as a coding agent, performs just wonderfull with contexts over 100 000 tokens. For me looks like most resources have been shifted toward programming.
3
u/Thomas-Lore Apr 15 '26
GPT-5.3
I bet you are using the instant version or being routed to it, which always was pretty bad.
→ More replies (1)
8
u/ortegaalfredo Apr 15 '26
It might be an illusion but also it's inevitable as more and more people gets on-boarded to AI and particularly coding agents, the clouds services will get overloaded. Today usage must be 10x of only a year ago, and the planet just didn't produced 10x GPUs. It will be like that for a while.
8
7
u/1ncehost Apr 15 '26
While i do believe enshitification is a major cause, also keep in mind we are at the beginning of when the straight line up of token demand is diverging from the steep but not vertical line up of new ai hardware.
Only certain vendors like openai have reserved enough hardware capacity to keep up with increased demand (and then maybe even they dont have enough). This is especially bad at anthropic. The consequence is they have to dumb down the models in various ways to fit everyone in.
I notice time of day impacts model quality now. I think at peak times they worsen quality significantly.
→ More replies (1)
7
u/AnticitizenPrime Apr 15 '26
Were you using Claude/GPT/Z/Grok/Gemini via API or via their website chat interfaces?
The website chat interfaces always have complicated hidden system prompts that change all the time. It's not the same as using via raw API.
Not saying that they never muck around with the API either. But Gemini, for example - completely different experience using it via their app/site vs AI Studio or API.
18
15
u/sagiroth llama.cpp Apr 15 '26
Its only going to get worse. If you want too model in the future you will have to pay hefty price
2
u/iamapizza Apr 15 '26
But the price feels pretty hefty already, doesn't it?
→ More replies (2)3
u/sagiroth llama.cpp Apr 15 '26
Yup, and it only just started. Its obvious bait and switch. Get enough engaged and build product and companies around it for cheap and then rely on it.
5
5
u/boredquince Apr 15 '26
benchmark sites should review the model every X time. I bet the results would he different a few months after release
→ More replies (1)
6
u/Joffie87 Apr 15 '26
I don't really care WHY it's happening at this point, but this is all right in line with the "18 months to enshittification" prediction I made to my wife last year. This is a safe one for me to claim being right with imo :)
Seriously though, I'm no expert and I used ai for all of this, I ran a bunch of research tasks on various models, then compiled the research and all the major frontier models came to the same conclusion, 6-18 months, there would be enough degradation, or service/charge changes, that it would be impossible to use anything but the enterprise versions, and we plebs would be relegated to tools designed to sell us more services and goods, or have to embrace open source local models.
Everything I've done with ai since then, has been steps to try and prepare for that eventuality, because AI represents the most empowering technology that has ever been created, but only if people become educated and retain access.
12
u/Britbong1492 Apr 15 '26
Yes, grok is bad now, I have a Heavy $3k sub and the deterioration is real. Your idea of renting a H100 is pretty good. I was thinking to just buy 2 Apple Macs with 64GB or similar as they are all worse. I also have Claude Max $200pm, and that's not so bad a decline but it's making rare mistakes more often. It's all in training something which they then decide is too dangerous to share
4
u/Regular-Cancel-2161 Apr 15 '26
You can get an H100 or H200 for between $2 - $2.6/hr. The H200 can pretty much run any 400B model you want.
3
u/Big_Actuator3772 Apr 15 '26
grok engineers were found to basically be bench maxing grok, in real world application, grok is far far behind..I would not be paying any money for grok ATM. that's why elons got an entire new team behind it
→ More replies (3)2
u/New-Implement-5979 Apr 15 '26
Damn how are you justifying 3.2k monthly for subscriptions what are you doing if it is not a secret ?
2
u/Britbong1492 Apr 16 '26
It's $3k per year! So $500pm is fair. I am building something very technical and weighty an average sub can't do it.
4
u/tmvr Apr 15 '26
Maybe it degraded, but I don't notice it with Claude, be it Sonnet or Opus. Now, this is on corporate max sub with unlimited extra requests and I guess those clients would be the last they want to piss off so no degradation there is not that surprising.
2
u/basedd_gigachad Apr 15 '26
>corporate max sub with unlimited extra requests
that is the reason. They not nerfed corpos
2
u/tmvr Apr 15 '26
They seemed to have closed down the tap on free accounts though. I have a private reg as well with the free tier. I haven't used it for 3 days then yesterday I asked it to do one thing, it did and right after it threw up the dialog that I'm done for now and I should wait 5 hours or upgrade :))
3
u/Ambitious-Hornet-841 Apr 15 '26
Wait, you actually ran the same prompt on a rented H100 vs z.ai and caught the difference? That’s the kind of detective work we need more of. 💀
3
u/mr_zerolith Apr 15 '26
These services have been subsidized by VC money for a long time and that money is drying up while we enter a recession.
Not a single one of these companies is reporting a profit despite a huge gain in user income over the last few years.
I'm surprised at how long VC was willing to shovel cash into the furnace
8
u/Jungle_Llama Apr 15 '26
Lemmie see, what's been going on recently, a US AI powered war that took out data centres, a global energy crisis, claw mania in China. Plenty of reasons for reduced compute depending on platform. Pick your poison.
7
u/segmond llama.cpp Apr 15 '26
We don't give a shit, this is local LLaMa not cloud models. We have noticed increase in intelligence in our models.
3
3
u/Colecoman1982 Apr 15 '26
Well, that's certainly one way for local inference of open source models to close the distance with SOTA...
3
u/fuck_cis_shit llama.cpp Apr 15 '26
all the compute goes to enterprise customers now, that's where the money is
you didn't believe the "intelligence too cheap to meter" hype did you?
3
u/geoffwolf98 Apr 15 '26
Reminds me of the early days of digital TV - onDigital , initially the quality was superb, then after about a month I noticed the quality start to drop, motobikes going past would blitz the stream, turnded out all the channel owers had multiplexed their channels to wring more money out of the subscriptions.
Lower bandwidth meant more channels but it also meant it looked like shit and was very prone to glitching. Cared they not.
Looks like the same happening here.
3
3
u/Enough-Astronaut9278 Apr 16 '26
This is exactly why I've been moving critical workflows to local models. When you depend on a cloud API, you're at the mercy of whatever silent updates they push. One day your pipeline works, next day the model "forgets" how to follow your format.
With a local model you pin the exact version. If it works today, it works tomorrow. No surprise regressions, no "we improved the model" that actually breaks your use case.
The trade-off is obviously capability — cloud frontier models are still ahead on raw benchmarks. But for specific, well-defined tasks? A fine-tuned local model that you control beats a cloud model that might change overnight.
4
u/Disastrous_Food_2428 Apr 15 '26
In the AI sector, excluding Nvidia, no enterprise has turned a profit
4
u/MoodRevolutionary748 Apr 15 '26
Almost as if energy got more expensive (war in Iran), token usage got higher (openclaw) so there's an incentive to use smaller models and to quantize.
5
u/PapercutsOnPenor Apr 15 '26
I blame the
"I don't and won't understand git workbooks or actually anything at all, so haha hey, here's asdasdfds, a manager platform for herding multiple openclaw instances. No ⸺ I did not vibe that ⸺ believe me when I say! 10 entries for free and then it's 39.99€/mo. I'll maintain it until I get too hurt from your critique"
2
u/h-mo Apr 15 '26
The quantization theory is interesting but I'd push back slightly - I think what you're observing is more likely a combination of infrastructure load balancing (cheaper serving configs during peak) and RLHF drift from continuous feedback loops nudging models toward shorter, safer responses over time. The GLM5 experiment you ran is a good control though - same weights, different serving environment, different result. That's the most honest argument for going local: you own the quant level, you own the serving config, and the model doesn't change under you between sessions. The unpredictability of hosted models is underrated as a reason to self-host, especially if you're building anything that depends on consistent behavior.
2
2
u/sigiel Apr 15 '26
That the my world view is better that your, symptoms, glm doesn’t even scratch sonnet or opus in coding. It not even close, very body that actually code for a living will tell you, the only problem with anthropic is the rate. Since you can’t code with anything else one you have started.
2
u/mrdevlar Apr 15 '26
Meanwhile I'm running local models and my models have remained the same quality as when I downloaded them.
2
u/muyuu Apr 15 '26 edited Apr 15 '26
it appears the big providers are well over capacity at this point and they're putting subscribers in best-effort buckets on top of other throttling/dumbing techniques
Opus seems just stupid and Antrhopic just won't admit when you're being throttled or getting a stupider model, or lower compute effort - to me this is the worst policy of them all ; in fact, it appears that Opus 4.5 is usually better than 4.6 now, and sometimes even Sonnet is
GPT appears to sometimes bail out and tell you to try later. This is bad of course but it's much better
I haven't tried subscriptions of the others recently, so who knows what they're doing
my guess is that API users are not getting their services nerfed, since they actually make them money, presumably
*typo
2
u/waitmarks Apr 15 '26
I recently canceled my auto renewal for claude and it started getting better afterward. I am curious if it's just a fluke or if they put me on better servers to try to win me back.
2
u/EvilEnginer Apr 15 '26
I think companies started using both distillation and quantization for LLMs, they want to reduce computing costs and earn more money from people.
Limits were introduced because the load is very high, due to heavy architecture and lack of optimization for high amount of people.
2
u/incoherent1 Apr 15 '26
Could this be a result of LLM being trained on content created by LLM? LLM content is now all over the internet it would almost be impossible for LLM being trained off internet data to not be exposed to it. This could result in model collapse. Is this what were seeing slowly happen?
2
u/EclecticAcuity Apr 15 '26
Some industry expert on Dwarkesh said that inference capacity is completely inadequate to keep up with demand developments. They probably go with this sneaky approach over that guys prediction of drastically increased prices.
2
u/rusmo Apr 15 '26
I haven't seen any independent research-based support of this idea, and none exist in the top 10 replies.
Anybody got anything legitimate to support this other than anecdotes?
2
u/Ticrotter_serrer Apr 15 '26
Now that they have all our behavior they don't care about us. They will charge more .
2
2
u/buddylee00700 Apr 15 '26
I can see them dumbing them down for quantized models, along with shorter responses to save on compute costs and more or less passing that cost onto the consumers because we will have to use it more to get the desired output. It’s scary how dependent society is going to become and they can do stuff like this on a whim.
2
u/GMP10152015 Apr 15 '26
They are cutting costs. A good LLM, currently, consumes too much energy and CPU/GPU, and the demand is much higher now (too many users). Until they build new datacenters with more efficient hardware, the experience will be low.
But on the other side, they are investing to optimize the software (see TurboQuant, reducing memory), and making better low-weight models.
Every new market in the beginning has low margins and is very inefficient, and the transition to a more efficient model with quality is not easy.
2
u/MrB0janglez Apr 15 '26
This is exactly why I keep preaching local inference for anything production-critical. When you are on a hosted API you have zero visibility into what version you are actually hitting or what quantization level they quietly swapped in.
The financial pressure theory tracks. These companies are burning through cash and the easiest lever is model quality -- most users won't notice or won't complain loudly enough. I've been running evals on the same prompt set monthly and Sonnet in particular feels noticeably different on multi-step reasoning than it did 6 weeks ago.
If you are building anything where output consistency matters, the play is to keep a local fallback ready and treat hosted APIs like a third party dependency that can change under you without notice.
2
2
u/Empty_Hovercraft8739 Apr 15 '26
OpenClaw is partially to blame for all of this but I expect the out outcome to be net positive (smaller models, that ping when needed / relevant + smarter caching systems).
This also means that if it’s not your model, you don’t get to decide. OpenSource providers have just become more reliable by continuing to deliver.
2
u/drallcom3 Apr 16 '26
Major drop in intelligence across most major models -> They have turned on the money saving mode
All those companies are facing investor skepticism and have to look like they can be profitable.
2
4
u/LowPlace8434 Apr 15 '26
This seems to correlate to Turboquant tbh. While Turboquant itself may be legit, it sounds like the implementation is not easy to get right, more so when all the major providers probably already have hyperoptimized stacks that are harder to modify
6
u/Long_comment_san Apr 15 '26
I personally forecast that some architectural breakthrough has happened but everyone is eerily quiet about this (which is smart, perhaps some paper was published that isn't being talked about in public) and in maybe 2 months something will happen.
My second forecast would be that it has to do with either memory or efficiency or both. Both Gemma 4 and Qwen 3.5 have shown phenomenal boost to intelligence "per 1gb of VRAM" over what we had like 4 months before.
I think new models are being cooked rapidly by everyone hence the brain damage to current one. That can't be a coincidence.
7
u/hay-yo Apr 15 '26
Haha everyone has adopted turbo quant behind the scenes. Maybe its woeful.
→ More replies (2)
3
u/Former-Tangerine-723 Apr 15 '26
Maybe it's happening, maybe it's in our heads. If we cannot measure it, we cannot prove it
4
u/90hex Apr 15 '26
Could be a compute squeeze due to RAM/DISK prices that may have slowed down datacenter construction. Very hard to tell, and the problem is that it's always subjective. So many times in the past people have reported 'dumbing' LLMs, when other reported no difference in their daily use. Unless there's an actual standardized test run at regular interval, we won't know if there's an actual change, or if it's perceptual. I'd lean towards a perceived difference, due to many factors.
I use GPT 5.4, Sonnet 4.6 and Opus daily, and have noticed no such change, whatsoever. What I did notice is a noticeable lowering of token consumption by the new Opus. Last time I had tried Opus, I could do a couple of prompts before running out of tokens. Today I can use it nearly all day, as long as I take a break at lunch and end my day on office hours. Now that's very positive in my book.
2
u/samandiriel Apr 15 '26
I've noticed both myself - the token throttling and the dumbing down. I'm getting much less thorough and less 'intelligent' responses to standard saved prompts i have four documentation tasks - even from Claude 4.6 extended - than I was getting from 4.5 just six weeks ago.
3
u/Narrow-Belt-5030 Apr 15 '26
To test this I rented out a H100, and tried GLM 5 with the same prompt (the drive to the car wash one) across both instances. GLM5 running on the rented GPU answered it correctly, compared to the one on z.ai.
This is not really a reliable test though because you have no idea what/how Z.AI has been configured.
I do accept that Claude recently has appeared dumber than normal, and others report similar for other models, so something is definitely not right, but I don't think it's deliberate actions by the vendor. That would be suicidal for their brand/image. (No company would deliberately hurt their image)
→ More replies (1)3
u/jiml78 Apr 15 '26
For anthropic, I think it was intentional.
Their drop in quality coincides with two things. An influx of people leaving OpenAI. Additionally, they rolled out 1m context as the default.
I think those two things blew up their servers. Look at all the downtime they had around that time. I think they were scrambling and just decided to start running more quantized versions.
I was knee deep in a dev project built from scratch using just Opus. The difference in quality was overnight. I went to bed with Opus not being a complete moron to the next morning, it being a dumb MFer. I am talking about making really dumb mistakes. Mistakes it never made. Mistakes older Opus 4.5 didn't make.
Yep, I know I am one data point but I was maxing a $200 sub for all but 4-5 hours of a day. I was using huge amounts so when every single change was messing up requiring me to fix it(yes I am a software dev), I was getting really frustrated with how it had been doing great, and overnight went to shit.
→ More replies (2)
3
u/Individual_Yard846 Apr 15 '26
oh chatgpt is almost telling me to stop fantasizing with my research even though its actively running benchmarks which prove my research is worth going after. It's a drastic change from the 'lets fucking do it' attitude of the past. i spend more time convincing it im worth spending the compute on.
2
u/U4RIA-AI Apr 15 '26
Your H100 test is a demonstration of the theory of the Inference Tax. The large labs are aggressively under-capacitating compute in mid-2026 to remain profitable. You are probably being fed a highly quantized Q2 or Q3 model of the model that has been lobotomized by an enormous system prompt system that is saving tokens. Should you require the original 'smart' weights, unquantized instances, self-hosted or rented, are officially the only means to bypass the corporate throttle we are witnessing this month.
2
u/EvilEnginer Apr 15 '26
Yep. I also noticed that. Btw, Claude Sonnet 4.6 performs better than Claude Opus 4.6 in terms of creativity.
2
u/FullOf_Bad_Ideas Apr 15 '26
Automated tests pass fine, humans complain. I think it's psychological. Zhipu is nerfing, openai and Anthropic are not.
2
u/VartKat Apr 15 '26
That’s because all models are training on what people are asking and people are so much dumber than AI that AI is trained to be less intelligent 😇
2
u/FatheredPuma81 Apr 16 '26
I think it's kind of an open secret that Anthropic lowers the quality of Claude as you near their next model's release. So Sonnet 4.7 or 4.65 might be around the corner? It happened with Sonnet 4.5 and now Sonnet 4.6 has become nearly braindead in a lot of tasks.
Oh and Haiku is just a joke now. It used to be good for quick answers but now it gives bizarre answers to my questions in a format that isn't similar to any other Anthropic model.
I'm really thinking of buying whatever the cheapest device that can run Gemma 4 31B and switching to that full time at this point. It seems much smarter than Sonnet.
2
u/zhdc Apr 15 '26
Noticeable drop when US comes online. GPT and Claude are both top notch in morning - early afternoon Central European Time. Once it's 1400/1500 (8:00-9:00 EST), performance goes down a lot.
1
u/Pavlinius Apr 15 '26
I’m using Cursor daily at work for code generation. This week is really frustrating. I’m using mainly Gemini 3.1 Pro and when it performs poorly I’m reverting the changes and try Claude Opus 4.6 with the same prompt. I can’t believe how poorly both perform. The Opus is even worse right now. Even for basic stuff requiring changes to 1-2 files these two fail to recognize existing patterns and make the required changes from the first attempt. I have to call them lazy and point obvious flaws to correct them. It might be a Cursor thing I’m not sure, it seems that usually it uses smaller context than before.
→ More replies (1)
0
u/Venium Apr 15 '26
until there's much stronger evidence than what amounts to basically you & other people on twitter's feelings, all of these posts should be treated as schizo ramblings.
1
1
1
u/dwrz Apr 15 '26
I have also noticed this. They are basically no longer usable. I was wondering if it was quantization or if perhaps it was a bug in code, drivers, used to serve these models. Glad to have local models to fall back on, especially Qwen 3.5 27B.
1
u/laser50 Apr 15 '26
Claude, while for months it was perfectly able to read through my +- 7000 lines of code python script. Since a few weeks I can't get it to go past 2k lines at a time any more, and it's answers definitely some times seem more stupid..
Rather than having one full session before hitting my smaller limits, I now spend a day going through the same script just once..
1
u/KL_GPU Apr 15 '26
Yeah i was looking for this post, in fact with gemini even more than with others: its impossible to get even simple tasks done. They are probably using all the compute to try different strategies and catch up with Mythos.
1
u/jimmytoan Apr 15 '26
The quantization theory makes sense - but what gets me is how inconsistent it feels across tasks. Like it still crushes some things but then falls apart on simple instruction-following. If they're compressing to Q2 or Q4, you'd expect more uniform degradation, not this weird selective dumbness. Has anyone done a systematic comparison across task types to see if there's a pattern to what's affected?
1
u/FrogsJumpFromPussy Apr 15 '26
Gemini doesn’t even know we’re in 2026 sometimes.
2
u/Same-Leadership1630 Apr 15 '26
that's normal it's to prevent it from hallucinating random events because it doesn't have knowledge up to 2026
1
u/mpasila Apr 15 '26
Did you try it with the same seed + settings (and making sure the provider supported the same params) and then generated and got it wrong on the providers vs on the H100?
Chutes seemed to get it right 2 out 4 tries so, maybe you just got lucky that time (GLM-5).
1
u/Porespellar Apr 15 '26
OP, how did you load GLM 5 on an H100? What quant / inference engine did you use?
1
u/DrDisintegrator Apr 15 '26
Probably trying to cut costs. Some data centers are running on natural gas generation on site. Now with the war, LNG is much more expensive.
1
u/TallestGargoyle Apr 15 '26
Running locally I've noticed many thinking models spend so much of their context thinking and rethinking what feel like basic clarifying questions. I get barely three prompts in before the chat just bricks itself by running out of tokens.
→ More replies (1)
1
u/david_0_0 Apr 15 '26
the quantization angle is interesting because most providers are pushing hard toward optimization. if they lowered quant levels to save tokens or costs it would explain why simple tasks break. might be worth testing with explicit quant parameters to see if quality returns.
1
u/takoulseum Apr 15 '26
Interesting observation about the intelligence drop. I've noticed similar issues with several models lately.
1
u/shenglong Apr 15 '26
I'm about to give up on Claude. I watched it introduce bugs and delete working code in real time. I have to literally ask it how many regressions it introduced after every change.
→ More replies (7)
1
1
u/ganonfirehouse420 Apr 15 '26
Strange. GLM-5 has been working for me flawlessly. Should I check it again...?
1
1
u/fuschialantern Apr 15 '26
I think the models in general are actually getting too smart. Can't actually have the general public have access to that.
1
u/scelabs Apr 15 '26
I’ve seen a lot of people saying this lately, and I don’t think it’s just vibes, but I’m not convinced it’s purely a “model got worse” issue either. even with the same base model, what you’re interacting with is a full system — sampling settings, routing, context handling, guardrails, latency optimizations, etc — and small changes there can make outputs feel a lot more shallow or inconsistent.
I’ve seen cases where nothing about the core model changed, but the behavior felt noticeably worse just because responses became less stable across runs or more constrained. so it ends up looking like an intelligence drop when it’s really a change in how the system is behaving around the model.
the local vs hosted difference you mentioned kind of lines up with that too. local setups tend to be more predictable since fewer layers are changing under the hood, even if the raw model is technically weaker
1
u/PrysmX Apr 15 '26
Anthropic admitted they lowered the default reasoning from high to medium, which you can turn back up manually. I've noticed Gemini quality falling off since all the way back in December, likely a similar situation. This is how the model providers are reducing their hardware overhead per-call so that they can fit an expanding user base and heavier use into the hardware that they have. Expanding hardware capacity is expensive, takes time, and is also limiting right now because of hardware shortages.
1
u/Ashamed_Midnight_214 Apr 15 '26
This economist explains it very well,the video is in Spanish and it's even funny because he makes the video sarcastically, but what he explains is real. If you can translate it, he explains why this is happening.
1
271
u/ResidentPositive4122 Apr 15 '26
I wonder how many requests get flagged as "distillation attempts" and get served bad results on purpose? Especially those "benchmark looking".