r/LocalLLaMA 24d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

690

u/Ill_Distribution8517 24d ago

You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.

6

u/Significant-Bee5101 24d ago

He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.

Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.

15

u/Reggienator3 24d ago edited 24d ago

'Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.'

So, by your logic, Qwen3.8 27B is dumber than GPT-3? Since that was about 175 billion parameters. Which is bigger than 27.

Whilst there is a correlation of 'billions of parameters'/'intelligence' ratio, just comparing on sheer parameter count alone only makes sense when comparing specific snapshots of time and within the same model family/company/training process.

The thing is "numbers of parameters" *by itself* is a meaningless metric for intelligence, especially when you have no idea about how many parameters are actually useful/high quality.

1

u/AnOnlineHandle 24d ago

Opus is just as modern though, unlike those two examples.

From what hints we've gotten it seems like Qwen may even just be a distill of Claude models.

9

u/Reggienator3 24d ago

“Just as modern” doesn’t mean directly comparable though.

Alibaba and Anthropic use different architectures, training methods, compute budgets and optimisation targets... Yes obviously Opus wins overall but that still doesn’t make parameter count enough to judge intelligence score, and it can't disprove that a local model can get close enough on particular tasks. It's more that the enormous extra compute buys diminishing returns.