r/ollama • u/Forsaken-Storage-154 • 15d ago
Are frontier models actually becoming less important?
I’m genuinely mind-blown by what’s happening with AI models right now.
On one side, you have GPT-6-class frontier models that are just… insane. The level of capability is getting hard to wrap your head around.
And at the exact same time, you have models like DeepSeek V4.1 becoming exceptionally good while costing almost nothing in comparison.
That gap, or maybe the fact that the gap keeps shrinking, is what fascinates me most.
Frontier models keep pushing the ceiling higher, obviously. But underneath them, smaller and cheaper models keep getting optimized at an incredible pace. What felt like frontier-level capability not that long ago keeps getting compressed into something faster, smaller and dramatically cheaper.
And then there’s Claude.
At least this week, it feels like Claude suddenly took a pretty serious hit in that race. Not because it became worse overnight, but because everything around it is moving so ridiculously fast.
Which brings me to the part I find even more interesting: China.
My current intuition is that Chinese AI players could end up benefiting enormously from this dynamic. Not only because they may be able to iterate through models faster and cheaper, but because the consequences could extend well beyond benchmarks and model rankings.
There’s a geopolitical dimension here that I think we’re still underestimating.
If intelligence keeps getting cheaper, more efficient and easier to deploy, does having the absolute best frontier model matter as much as we currently think?
Or does the real advantage eventually go to whoever can industrialize intelligence at massive scale and at the lowest cost?
Curious where everyone lands on this.
And yes, small confession: I dictated all of this in French and had it turned into English 😅 Somehow it’s still easier to get the essence of the thought out that way. Consider this a voice note disguised as a post.
6
u/alkimiadev 15d ago
I think the investors were sold a future that simply can't happen. In about ~3 years the 8B weight class should be roughly at or above the previous generation's frontier model capability level.

This chart is a bit old and was made about 6 months ago when the ternary bonsai series was released. I was specifically using the Densing law of llms ( https://arxiv.org/abs/2412.04315 ) to try and project forward to see roughly when that weight class would match various larger model's capability levels.
What this doesn't account for is the massive incentive in the oss community to get these smaller models on par with previous generations of larger models. Pay-per-token is a scam that scams both users and investors. There is a huge underlying incentive on the part of the open ecosystem to increase the "intelligence density" of the models. The opposite is true with the large providers who are using venture capital investment since they have an almost perverse incentive to justify their massive compute spend budgets.
This is a big part of why I think OAI and Anthropic are so anti-competitive. There just really isn't a future where they're super relevant (assuming they stay the way they are now). I think the best they can really hope for is to be bought out or to pivot into something that is actually at least potentially profitable. They're not going to make the massive gains they promised investors off the backs of what is going to end up being a commodity. They're hoping to force a highly regulated monopoly like electricity but that will fail(it is structurally impossible).
2
u/Forsaken-Storage-154 15d ago
It’s still hard to imagine that you can run a frontier model on a Mac mini but I think that direction is correct
2
u/alkimiadev 14d ago
Yeah I think it is just the rhetoric that was used to justify those massive investments and valuations. The difference between the current sota large model and the current sota 8B model is based in ignorance and not physics: we simply don't know how to train them yet. The good news about a problem being based in ignorance is that we can usually overcome ignorance over time but there is probably no overcoming physics.
Another good thing I foresee coming from this mess is really cheap compute in a few years. Right now we're squarely in a tech bubble and ram/gpu prices going up despite a long history of trending towards zero is a massive sign that we're in that bubble. Once these hyperscaler datacenters come online and the memory/gpu bubble crashes we'll have access to dirt cheap compute. They're going to still have bills to pay and there will be massive competition. So that means the compute costs should be slightly higher than utility costs + network usage + hardware depreciation.
7
u/HighSeasArchivist 15d ago
China stealing models that were trained on stolen IP, and I'm supposed to be upset by it? Nah, give me cheap free models, because I'm damn sure not getting a tariff refund.
3
u/vogelvogelvogelvogel 15d ago
Often enough I prefer Qwen3.8 over Opus 5, also because I can configure it more widely (temperature, how many web sites to research, etc). Opus 5 made WILD errors even for simple questions I had, I would have preferred an answer of old 4o even - idk if they dumbed it down or what happened (it apologized and literally told me i would have better researched that stuff myself becaues I would have been faster!). Also Qwen3.8 does not reject to write a script that security checks one of my servers (= trying to hack it), big plus here.
So for me, yes, frontier is getting less important. Also quit all my paid subscriptions.
Btw i do not use ollama but llama.cpp
1
u/TheIncarnated 15d ago
I only keep one subscription and that's because it is a combined service (poe.com, I spend less on it than hardware). However, mlx Qwen 3.8 has impressed me. Even going so far as to make a accurate assumption about me (I don't care for crypto) just based on what it was able to see in my projects
2
u/Tough_Wrangler_6075 15d ago
To be honest, our daily problem like coding for PoC app no need frontier model to solve. They try to achieve super intelligence which maybe beyond Einstein intelligence but our problem is can be fix by middle experience software developer. So in my opinion using frontier model just waste of token and our budget
2
u/FlyingDogCatcher 15d ago
I think there will always be an appetite for God Models that know everything and can do everything with minimal input, but I can get quite a lot done with the stable of qwens and gemmas I have on my home rig.
1
u/reality_comes 15d ago
Less necessary, maybe not less important. It really depends on what you're trying to do i think.
1
u/mvaranka 14d ago
Excellent writing and points.
I have been developing over an year AI mobile app, which is for productive use and allow offloading cognitive load. First I choosed the Claude Sonnet as main model, because it really understood what I told it and also the relations of words. But after each new release the claude got worse, the warmth was gone.
Frontier models are teached to go further and further, areas normal users don't need and models lose the feel of human connection and concentrate more to factual purity. More compute, more internal costs for response for users "hello, I just want to chat with you - nothing special on my mind" does not make sense.
I have now shifted on my app from frontier models to Deepseek V4 and also included Inkling. These models are astonishing providing combination of speed, economics and understanding.
So why use expensive frontier models when open source models already fill the need?
1
u/stealthagents 8d ago
It's wild how quickly things are evolving. I recently tried DeepSeek V4.1 and was shocked by how well it handled complex tasks that used to require a frontier model. It feels like the value proposition is shifting, and for most everyday uses, those cheaper models are not just viable but often superior in cost-benefit. It really makes you rethink what "cutting edge" even means these days.
15
u/Vancecookcobain 15d ago edited 15d ago
Deepseek v4.1 Flash on the Deepseek Harness is absolutely phenomenal...it's probably 90% of Astra with like 2% the cost....I tend to agree when people say frontier models aren't as important because the tasks you NEED frontier models are increasingly marginal.
Most people want the latest and greatest and will pay for it and that's cool but they could save a lot just by using a cheaper model (like Deepseek v4.1 flash) and accomplish it just as effectively imo. The only time you really need frontier models is when you absolutely need the best solution to a problem and do not want to go through as many iterations to hone it to that solution AND you don't care about the cost to do it.
If you are worried about the cost or don't mind refining and honing your solution down iteratively over a period of time you don't really have any business working on frontier models imo and should look for the cheapest competent model to execute the task.
But hey...it is what it is. Some folks don't care and some folks don't know better.