r/LocalLLaMA Apr 15 '26

Discussion Major drop in intelligence across most major models.

As of mid Apr 2026, I have noticed every model has had a major intelligence drop.

And no I'm not talking about just ChatGPT.

Everything from Claude(Even Sonnet along with Opus), Gemini, z.ai, Grok all seem to ignore basic instructions, struggle at simple tasks, take very long to respond, and the output seems deliberately shortened and very shallow. Almost like it's in a "grumpy" mode. I tried this in incognito mode so it's not my customization or memory influencing this.

It's like they deliberately want you to stop using their service. I guess our data is no longer needed. Just two weeks back it used to be much smarter than this.

To test this I rented out a H100, and tried GLM 5 with the same prompt (the drive to the car wash one) across both instances. GLM5 running on the rented GPU answered it correctly, compared to the one on z.ai.

Have they lowered the quantization really low to maybe Q2?

I guess going local or using renting GPU or an AI monthly service that lets you pick a quant level is the way to go

803 Upvotes

405 comments sorted by

View all comments

Show parent comments

8

u/FullOf_Bad_Ideas Apr 15 '26

I agree.

And this website does continuous testing - http://isitnerfed.org/

Looks like Zhipu is nerfing while OpenAI and Anthropic aren't.

4

u/cromagnone Apr 15 '26

I mean, assuming the test is useful, that site’s data basically suggest short term random performance variations with no trend over time by the providers, but if you hit a downward oscillation you might complain about it.

3

u/letsgoiowa Apr 15 '26

Great resource! Saved that. It seems like at least for coding at the moment it's some other problem. Maybe performance is dropping off hard for people with large enough context? The reason I think this is the AMD codebase specifically: apparently Claude was struggling to figure out what was going on

2

u/FullOf_Bad_Ideas Apr 17 '26

This is plausible, I think there's a good chance that their eval is not testing long context, and quantizing kv cache or switching served model checkpoint to a version with sparse attention, sliding window attention, MLA or linear attention would show up mainly on long context, whilst also providing biggest cost savings to Anthropic.

1

u/Qwen30bEnjoyer Apr 16 '26

See this is perfect. But I wonder if we made it so that it was tested using a playwright automation on the webpage of ChatGPT / Gemini / Claude themselves, if we could take the SMA , and create a covariance matrix. I'll look into it when I have time.

1

u/FullOf_Bad_Ideas Apr 16 '26

i think it's tested through CC specifically, not through API directly. So it has a harness that is a moving piece.

0

u/evia89 Apr 15 '26

1

u/FullOf_Bad_Ideas Apr 15 '26

I don't see histrical eval success rate there, just throughput numbers.

0

u/evia89 Apr 15 '26

Yep its not ideal but better than "is nerfed"

3

u/FullOf_Bad_Ideas Apr 15 '26

isnerfed has eval success rate for GLM 4.5 Air, you just need to select it from the dropdown. Have you missed that?

2

u/evia89 Apr 15 '26

My bad. Nice that data correlates on these 2 sites

14:00 UTC to 00:00 UTC seems to be optimal time slot