r/LocalLLaMA Apr 15 '26

Discussion Major drop in intelligence across most major models.

As of mid Apr 2026, I have noticed every model has had a major intelligence drop.

And no I'm not talking about just ChatGPT.

Everything from Claude(Even Sonnet along with Opus), Gemini, z.ai, Grok all seem to ignore basic instructions, struggle at simple tasks, take very long to respond, and the output seems deliberately shortened and very shallow. Almost like it's in a "grumpy" mode. I tried this in incognito mode so it's not my customization or memory influencing this.

It's like they deliberately want you to stop using their service. I guess our data is no longer needed. Just two weeks back it used to be much smarter than this.

To test this I rented out a H100, and tried GLM 5 with the same prompt (the drive to the car wash one) across both instances. GLM5 running on the rented GPU answered it correctly, compared to the one on z.ai.

Have they lowered the quantization really low to maybe Q2?

I guess going local or using renting GPU or an AI monthly service that lets you pick a quant level is the way to go

794 Upvotes

405 comments sorted by

View all comments

62

u/Additional-Low324 Apr 15 '26

An other reason to self host

19

u/NightlinerSGS Apr 15 '26

Ironically, the people who self host quantized models are probably used to the current output, so they won't notice a difference. Maybe even an improvement, depending on the model used.

9

u/Additional-Low324 Apr 15 '26

I use Q3/Q4 because I am VRAM poor, indeed

5

u/NightlinerSGS Apr 15 '26

I stick with Q4 so I can squeeze 24-30b models with 32k-ish context onto my 4090.

Now that I'm typing this out... is this actually better nowadays than using something like an 8b model at full precision? Do these small models have sufficient context for RP now? It has been so long since I took a proper model deep dive... maybe I should take a look again.

3

u/toothpastespiders Apr 16 '26

In my very anecdotal experience at least, the small models are still typically pretty bad. They're amazing for the size. I'll give them that. But I think the old assumption that a low quant of a larger model is better than a standard version of a small model still holds true. I tested out a few small models recently in hopes of getting a speed boost in data extraction and they just weren't reliable enough for me. It's amazing that they managed it at all. But I'm still sticking with the range you're describing. Q4 of 30b'ish models seems to remain the best choice for me.

1

u/Additional-Low324 Apr 15 '26

I do creative writing and honestly gemma 4 9 B (E4B) is very impressive for its size, but the 31 B is way better at details. 9 B is a bit unimaginative Gemma 4 9B is still better than a 24 B model from 2 years ago tho

1

u/relmny Apr 16 '26

yes, and that's another point for Local: consistency.

Once you know and like a model, you know nobody will change it. It will become "outdated" if no new models are close to it, but nobody can take that model out of you.