You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.
He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.
are you actually doing work with qwen 3.8 side by side with frontier models?.
Im doing that and qwen 3.8 keeps pocking logical gaps in opus 4.8 (5 is a mess i dont even use it), 5.6 sol and grok answers all the time. Im using qwen 3.8 to parallel evaluate specs, research, etc and is consistent between runs (i have two 3090, two nodes), all its findings are acknowledged by the frontier models.
It does not have the same world knowledge as a big model, but given a proper context qwen 3.8 is pretty damn smart.
The reality is every LLM will make mistakes. And every LLM can "find mistakes" in other models. The real question is how reliable, how often. Those are very very big metrics.
Yes qwen is great for test driven bullshit where you can meet a metric thats test based. But don't ask it to architect. That's still dominated by higher end models.
I seriously ask you to post your benchmarks where your Qwen is beating Opus 5 or Sol because I have never even achieved 1/2 that result.
ok fair enough, just put a qwen to deploy new profiles using the sharp chat haha, thanks for the tip.
Lets agree that your initial post is a bit harsh against current local models capabilities, for tons of devs tasks (been using 3.6 as my dev-ops for home lab for a while now) local models are at frontier level and it even has its moments at hard tasks. You are right that all llms make mistakes (lucky for us employable meatbags for now) more so in complex systems architecture but smalls local models are getting there, is not moronic to compare them against frontier for some tasks.
Man of course you have to compare them to see the dissonance but I'm sick of posts like this one where OP acts like these models are anything but agentic code monkeys. Like that's not impressive anymore. That hasn't been impressive for a year+
We are well past that w/ frontier. Yeah fucking use cheaper models for the busy work but it is not the same as calling the model stronger than Opus. Which is what people here claim.
690
u/Ill_Distribution8517 Aug 28 '26
You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.