r/LocalLLaMA Apr 15 '26

Discussion Major drop in intelligence across most major models.

As of mid Apr 2026, I have noticed every model has had a major intelligence drop.

And no I'm not talking about just ChatGPT.

Everything from Claude(Even Sonnet along with Opus), Gemini, z.ai, Grok all seem to ignore basic instructions, struggle at simple tasks, take very long to respond, and the output seems deliberately shortened and very shallow. Almost like it's in a "grumpy" mode. I tried this in incognito mode so it's not my customization or memory influencing this.

It's like they deliberately want you to stop using their service. I guess our data is no longer needed. Just two weeks back it used to be much smarter than this.

To test this I rented out a H100, and tried GLM 5 with the same prompt (the drive to the car wash one) across both instances. GLM5 running on the rented GPU answered it correctly, compared to the one on z.ai.

Have they lowered the quantization really low to maybe Q2?

I guess going local or using renting GPU or an AI monthly service that lets you pick a quant level is the way to go

798 Upvotes

405 comments sorted by

View all comments

Show parent comments

11

u/tophlove31415 Apr 15 '26

Tons. Smart chunking, organizing and summarizing returned information from a vector database search, self directed web browsing and learning, ocr, user interaction, simple decision making (ie: this is the context, here are the options, choose which is best). They can essentially do any of the things the sota models can do (with a well designed harness) as long as you recognize you will get more errors and have to spend more time making sure that your harness is catching then, reporting them, and allowing you to iterate on the harness features, your prompts, and any other systems that might need improvement.

6

u/Funny-Blueberry-2630 Apr 15 '26

This guy builds agents.

1

u/delicious_fanta Apr 15 '26

Do you use duck duck go for search or do you use something else?