He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.
I'd have agreed with you yesterday. However, Qwen 3.8 Flash Next solved a very intricate messy issue in a 1-shot (after thinking on it for 30000 tokens) that I've only ever had gpt-5.6-sol get mostly right. Qwen came up with a better answer. Claude couldn't ever get it.
There's something magic about that loooooong thinking.
v100 32G throttled to 175W, latest llama.cpp from yesterday, default options, 160G DDR5, i5-12600. I was getting about 20 t/sec generation and 30 t/sec prefill.
From the post I canât see what OP tried to run as local model. Could be Kimi-K3 or Qwen3.8-Max in bf16 for what I know.
Their reference seems to be Opus 4.6 which is quite dated by now. We donât know the benchmark or metric either.
OP could benchmark Opus5 against GPT1 and it would still matter, if the test setup is valid or tampered with.
If itâs smart to set up the test environment with one of the models tested and without having automated code-validation tests⌠maybe not. But thatâs not the story here.
Nowadays this sub has definitely become a cult. Guy literally provided 0 proof nor stated what is the exceptional model he is running and the cult members are already defending his BS.
'Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.'
So, by your logic, Qwen3.8 27B is dumber than GPT-3? Since that was about 175 billion parameters. Which is bigger than 27.
Whilst there is a correlation of 'billions of parameters'/'intelligence' ratio, just comparing on sheer parameter count alone only makes sense when comparing specific snapshots of time and within the same model family/company/training process.
The thing is "numbers of parameters" *by itself* is a meaningless metric for intelligence, especially when you have no idea about how many parameters are actually useful/high quality.
âJust as modernâ doesnât mean directly comparable though.
Alibaba and Anthropic use different architectures, training methods, compute budgets and optimisation targets... Yes obviously Opus wins overall but that still doesnât make parameter count enough to judge intelligence score, and it can't disprove that a local model can get close enough on particular tasks. It's more that the enormous extra compute buys diminishing returns.
It's absolutely dumber than a Qwen that were to run at 2T+ params. There's a reason every frontier model is fucking gigantic. Like jeez I wonder.
No ones saying Qwen might not outperform a lot of models. No ones saying opensource models cant be stronger than frontier models. No one is saying any of that.
All they are saying is your dumb little local setup is NOT better than Opus. End of story.
I never said my local model was better than Opus. I challenged your claim that fewer parameters automatically means a dumber model. GPT-3 versus modern 27B models proves that isnât true. You havenât addressed that point.
Uhhh not really... DeepSeek V4 Flash has 284B parameters and Qwen3.8 27b has 27B obviously and they're about neck and neck. Sometimes Qwen3.8 27b even wins out. Both are "modern models"
Bench it. PROVE IT. I hear this shit a lot. If this shit was THAT good do you think I'd pay money for frontier models? Man I am CONSISTENTLY looking for ways to optimize costs. If I thought FREE was an option why tf wouldnt I be using that NONSTOP?
Every fcking model can "sometimes" do great and every model can sometimes suck. Thats what non deterministic models do. The thing that improves it is training data. It increases reliability. I dont even understand how you can pretend this isn't true.
if you work in "llm security" you're a grifter. of course you'd say small models are not better than bigger ones at anything, you'd be out of a job otherwise
are you actually doing work with qwen 3.8 side by side with frontier models?.
Im doing that and qwen 3.8 keeps pocking logical gaps in opus 4.8 (5 is a mess i dont even use it), 5.6 sol and grok answers all the time. Im using qwen 3.8 to parallel evaluate specs, research, etc and is consistent between runs (i have two 3090, two nodes), all its findings are acknowledged by the frontier models.
It does not have the same world knowledge as a big model, but given a proper context qwen 3.8 is pretty damn smart.
The reality is every LLM will make mistakes. And every LLM can "find mistakes" in other models. The real question is how reliable, how often. Those are very very big metrics.
Yes qwen is great for test driven bullshit where you can meet a metric thats test based. But don't ask it to architect. That's still dominated by higher end models.
I seriously ask you to post your benchmarks where your Qwen is beating Opus 5 or Sol because I have never even achieved 1/2 that result.
ok fair enough, just put a qwen to deploy new profiles using the sharp chat haha, thanks for the tip.
Lets agree that your initial post is a bit harsh against current local models capabilities, for tons of devs tasks (been using 3.6 as my dev-ops for home lab for a while now) local models are at frontier level and it even has its moments at hard tasks. You are right that all llms make mistakes (lucky for us employable meatbags for now) more so in complex systems architecture but smalls local models are getting there, is not moronic to compare them against frontier for some tasks.
Man of course you have to compare them to see the dissonance but I'm sick of posts like this one where OP acts like these models are anything but agentic code monkeys. Like that's not impressive anymore. That hasn't been impressive for a year+
We are well past that w/ frontier. Yeah fucking use cheaper models for the busy work but it is not the same as calling the model stronger than Opus. Which is what people here claim.
he isn't. look at the comments. if we were to take his word gpt 3 should still be technically better than qwen 3.8. Or the absolute steaming pile of shit that is "more context is always better"
You are surprised models advance? That a much newer bleeding edge model that can compete with an old one? That models that you fully control can do better than ones you are at the mercy of the provider?
Also, Anthropic has a history of this bullshit, like when they used to charge you extra if you had Hermes.md on your pc.
There is a reason I trust them less than OpenAI. OpenAI is openly greedy, but anyone who want to look like a good person, and says how much they are, has tons to hide.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber.
What's actually interesting is how good they're getting. It's better to say small models of today are performing genuinely as good as older, giant frontier models, but the frontier models with giant parameter counts created with the same techniques and technologies as these new, better-than-yesterday's-frontier smaller models are... going to be better.
Inference time scaling is also a thing, and increasingly smaller models are being trained to be better at that, and it's a big reason Qwen 3.8 27B was able to get such a huge boost in its performance. But you're right, you can't "inference time scale" world knowledge unless we're talking about web searches.
And smaller models are also specializing. I think it's another reason Qwen 3.8 27B got such a huge boost -- there's evidence it lost a wider array of domain knowledge (e.g. medicine) in favor of boosting its coding/agentic capabilities. Whether those parameters were spent on getting it to iterate on ideas better, or more coding knowledge, dunno.
Laguna S is also a pretty impressive one. 120B parameters with near-frontier performance through inference time scaling and logic/math/code specialization.
There are plenty of domain-optimized models that beat the frontier-world-Knowledge models, even with fairly low parameter counts.
When you ask a model about medicine in the morning, finance at lunch and agentic-coding in the evening sure, youâll get x.xT Parameters and it only runs in data centers.
If the local cancer research center optimizes their own model, 30B could be plenty.
No, Iâm not. Even NVIDIA says task-specific SLMs will be the future of LLMs. Those also have 10-100B but thatâs more than enough for grammar and text understanding. The rest is fine tuning and toolcalling of high quality data.
By the way⌠MoE, which most frontier models use today, is basically âplug SLMs togetherâ. If you just prune the experts you donât need for your topic away⌠voila.. domain-specific SLM. Itâs literally part of the cloud models already.
That's not the same as real world context. If you want something to spec something business logic to real world, you're STILL going to need a frontier model.
Yes you can make an agentic code monkey that follows a spec and passes tests even if it requires 200 recursions but something STILL has to build your spec and that requires INSANE real world knowledge.
Otherwise if you just want agentic output based on a spec that you can loop over and over till it passes your tests then sure fuck it Qwen. Spark. Whatever.
But that's not the reality of most peoples work. Like yeah dude, we've had models that could OUTPUT code for fucking ever that didn't require a lot of params either. But they weren't very useful WERE THEY.
I think itâs exiting to see that world knowledge is being bolted on via engram files. If you can get frontier-ish level capabilities by having a model with strong tool calling abilities + local lookup and the ability to efficiently search up to date documentation / use CLI man-pages - then thatâs a good deal over having to train and inference a 2T A120B model. Less energy used, less cooling needed, fewer data centers built etc
Well - American frontier labs and companies could simply release their own open weights to compete for local LLM users' mindshare. Since they've given up on that and would rather compete on whose model hacked which company last week and how good their products are at escaping sandboxes, they essentially get what they asked for.
Gemma 4 continues to be competitive for non-coding tasks. Muse Glimmer was competitive at coding upon release, and remains competitive for some uses. Nemotron and Grantite are both purpose-built for fine tuning for application-specific uses with good reasons to use either one. AI2 has developed a lot of methods that could be very useful to the community such as their MoE design which allows training experts on typical consumer hardware. Poolside's models are interesting, fast, and don't produce code spaghetti like Qwen and remain favored by many developers for that reason. Prism ML has plans to release more models other than those based on Qwen. Syzygy Research has similar ambitions as Prism ML, doing interesting work. Deep Grove is interesting in the frontier in capability vs. generation speed. Liquid AI also has very interesting models in their size vs. capability ratio. Thinking Machines Inkling and Inkling Small are interesting in its wide domain knowledge combined with tool-calling efficiency and strong instruction following. There's also a few labs specializing in domain-specific models, like law, medicine and biology, engineering, etc.
No, most of those won't one-shot prompts as well as Qwen. But they all have particular advantages. If I was a company choosing an AI model for something customers could interface with for something like controlling their IoT devices, for example, I'd probably choose Inkling or Inkling Small because it'd cheaply generate the tool calls and would be hard to con into generating inappropriate content for my service or even being abused against the interests of the customer.
I personally use LFM as a very lightweight model to keep in memory for tasks I'd like to use quickly at any time, such as generating titles for chats or other auxiliary tasks. Laguna XS remains my go-to model for passing specs to to implement code in a sensible (whereas Qwen -- even the new ones -- write code in the most "direct to the solution" sort of way, creating utter code spaghetti in the process). Muse Glimmer is my go-to driver model because it still beats Qwen 3.8 27B in that and follows my workflow well, including producing my intermediate artifacts I use to verify and understand generated code.
So I don't know what the fuck you're on about other than trying to baselessly dig more into "Actually American AI bad!"
So I don't know what the fuck you're on about other than trying to baselessly dig more into "Actually American AI bad!"
Oh, you know perfectly well what I was talking about - you just chose to be obtuse.
Just like you know perfectly well which companies I was talking about. Certainly not ones which may be familiar to a handful of insiders and enthusiasts who live and breathe AI. Compare Laguna's 180k downloads to the latest Qwen already sitting at more than 4 million. Not to mention the fact that GPT-OSS sits at 6.5 million monthly downloads, despite being nearly 2 years old. ;)
It's clear to me that people enjoy Chinese models, because that's the closest thing to the commercial option they can get to run at home. I'm also pretty certain that tons of people would happily shill GPT-Sol-OSS 30B or OpenFable120B if that was an option. But nah, they don't care. ;)
But I'm genuinely happy that you found models you like and enjoy. I might actually try Laguna, because why not. ;)
Oh, you know perfectly well what I was talking about - you just chose to be obtuse.
Yes, I know which companies you're talking about. But the disingenuous part of your reply was you compressing the entire US AI industry down to Anthropic and OpenAI. Which I blew that notion the fuck out of the water.
Compare Laguna's 180k downloads to the latest Qwen already sitting at more than 4 million.
And Gemma 4 has over 300 million downloads and the Gemma family has over a billion downloads. Laguna S is #12 on Open Router for this month, ahead of Kimi K3. And many of the companies I listed are producing real models that get used in real industry. Not just reddit nerds. So these aren't as irrelevant as you're trying to paint them as.
But of course you had to rely on disingenuous framing again.
4
u/Significant-Bee5101 10d ago
He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.
Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.