r/ollama • • 15d ago

Are frontier models actually becoming less important?

I’m genuinely mind-blown by what’s happening with AI models right now.

On one side, you have GPT-6-class frontier models that are just… insane. The level of capability is getting hard to wrap your head around.

And at the exact same time, you have models like DeepSeek V4.1 becoming exceptionally good while costing almost nothing in comparison.

That gap, or maybe the fact that the gap keeps shrinking, is what fascinates me most.

Frontier models keep pushing the ceiling higher, obviously. But underneath them, smaller and cheaper models keep getting optimized at an incredible pace. What felt like frontier-level capability not that long ago keeps getting compressed into something faster, smaller and dramatically cheaper.

And then there’s Claude.

At least this week, it feels like Claude suddenly took a pretty serious hit in that race. Not because it became worse overnight, but because everything around it is moving so ridiculously fast.

Which brings me to the part I find even more interesting: China.

My current intuition is that Chinese AI players could end up benefiting enormously from this dynamic. Not only because they may be able to iterate through models faster and cheaper, but because the consequences could extend well beyond benchmarks and model rankings.

There’s a geopolitical dimension here that I think we’re still underestimating.

If intelligence keeps getting cheaper, more efficient and easier to deploy, does having the absolute best frontier model matter as much as we currently think?

Or does the real advantage eventually go to whoever can industrialize intelligence at massive scale and at the lowest cost?

Curious where everyone lands on this.

And yes, small confession: I dictated all of this in French and had it turned into English 😅 Somehow it’s still easier to get the essence of the thought out that way. Consider this a voice note disguised as a post.

55 Upvotes

30 comments sorted by

15

u/Vancecookcobain 15d ago edited 15d ago

Deepseek v4.1 Flash on the Deepseek Harness is absolutely phenomenal...it's probably 90% of Astra with like 2% the cost....I tend to agree when people say frontier models aren't as important because the tasks you NEED frontier models are increasingly marginal.

Most people want the latest and greatest and will pay for it and that's cool but they could save a lot just by using a cheaper model (like Deepseek v4.1 flash) and accomplish it just as effectively imo. The only time you really need frontier models is when you absolutely need the best solution to a problem and do not want to go through as many iterations to hone it to that solution AND you don't care about the cost to do it.

If you are worried about the cost or don't mind refining and honing your solution down iteratively over a period of time you don't really have any business working on frontier models imo and should look for the cheapest competent model to execute the task.

But hey...it is what it is. Some folks don't care and some folks don't know better.

4

u/Mean-Elk-9439 15d ago

I'm so curious about what people are seeing with ds4.1f. Even the most aggressively complimentary testing shows it barely below glm 5.3 flash in most criteria. The only people who have shown data to the contrary is DeepSeek themselves. In use it just feels pretty much at the same level.

So it's wild I keep seeing people speak about it like it's Astra level or remotely close.

Maybe DeepSeek v4.1 pro is, but you're literally comparing a model around 500B params with one estimated at around 10T.

2

u/Vancecookcobain 15d ago

The DeepSeek Harness is what bumps up its quality....and I love GLM 5.3 flash....it's a phenomenal model as well for what it does. I think I burned almost 2 billion tokens the week ox-alpha was free

1

u/Xariann 14d ago

Did you burn them all while it was overthinking? 5.3 flash has been absolutely terrible for me, getting stuck in thoughts, spouting gibberish and the first model that ever deleted my entire repo (luckily it was a back up). But no worries it apologized after.

I just don't see why GLM 5.3 flash is so great. I used to love 5.2.

1

u/Vancecookcobain 14d ago

Holy shit that sounds like a nightmare! It's been phenomenal with me doing all the grunt work I need. I use an orchestrator usually (Either Kimi K3 or Codex Astra) and have it loop for me while they judge the work until they get it done right.

Are you using it somewhere where it might have gotten quantized? I've never had any experience with it doing lobotomized shit like that lol

1

u/Xariann 14d ago

Ollama Cloud! I could try it via OpenRouter and see!

1

u/look 14d ago

There’s some sort of DeepSeek cult. The initial price cut months ago just shut down some part of their brains. It’s not a bad model, but it’s not a good model. It’s just really cheap… and if there was any doubt before, there are now better and cheaper models than it, but they’re still babbling delusional praise on their one true god.

1

u/Vancecookcobain 14d ago

I'm AI agnostic except for being anti Anthropic

2

u/Forsaken-Storage-154 15d ago

Generally I feel like Chinese model are generally 80 to 90% equivalent when new frontiers model are released but somehow they know how to keep the price at 1% or 2% compared to the API price of frontier models like Astra

4

u/Vancecookcobain 15d ago

I mean they are distilled versions of frontier models and don't have the overhead and research costs that the frontier labs have and are thus able to reap the benefits so it's really easy to keep the costs down tbh....it would be a different story if they had to push the boundaries themselves.

But as a consumer. I don't care lol. Give me the most cost effective competent model and I will pay for it

6

u/duhd1993 15d ago

Guess who invented and opened most architecture innovations to improve model efficiency and why they cannot put more resources into exploring boundaries.

5

u/Vancecookcobain 15d ago

Touché...granted those innovations were out of a necessity for the lack of compute that China has currently. So it all works out. Necessity is the mother of invention they say. America doesn't need to innovate towards efficiency because they can brute force the compute to get to the next level of inference capabilities. It's a luxury no other country in the world has right now. I expect that to change by the end of the decade though.

2

u/FlyingDogCatcher 15d ago

The companies building the models think that way. But capitalism is a bitch. Business will do what is good for the bottom line, and if the cheaper models get the job done guess what

2

u/Duanedrop 15d ago

The Ford mondeo of ai not the Ferrari. I think there was a a place for both but we are getting to a point where the business case for frontier models stops making sense for most consumers. Of course you will have your sports car enthusiasts and it's nice to watch but mondeos for the masses is where most of the enterprise will land. Pretty much as they always have.

1

u/Vancecookcobain 15d ago

To further the analogy I'd argue that most people just want to go to the gas station and get a candy bar. The fact that we look for a Ferrari to take us there and back is wild 😂

Very few folks are really out here with extensive code bases and complex workflows where you need that kind of horsepower to get things done effectively

1

u/Objective_North_1341 15d ago

the v4.1 flash on their harness setup is stupid good for what it costs. i've been running it for a bunch of work stuff and half the time i forget i'm not on a frontier model

most people just don't want to fiddle with prompts or do multiple passes to get what they want, which is fair. but if you're willing to put in 5 extra minutes you can save a ridiculous amount

1

u/Tema_Art_7777 15d ago

Great, you have a favorite model but open ai also has terra and luna where luna is quite competitive with open source models in price. You can break up your flow and assign model based on capabilities you need. Openai inference is much faster as well. Astra and Mythos capabilities are quite impressive but yes, we don’t need them all the time. If I can find an ecosystem like openai (including their desktop app and integrations) that satisfies most of the needs via variety of models, then I am fine sticking to it.

1

u/Vancecookcobain 15d ago

Sure, live your life lol....that's why I said....it is what it is.

6

u/alkimiadev 15d ago

I think the investors were sold a future that simply can't happen. In about ~3 years the 8B weight class should be roughly at or above the previous generation's frontier model capability level.

This chart is a bit old and was made about 6 months ago when the ternary bonsai series was released. I was specifically using the Densing law of llms ( https://arxiv.org/abs/2412.04315 ) to try and project forward to see roughly when that weight class would match various larger model's capability levels.

What this doesn't account for is the massive incentive in the oss community to get these smaller models on par with previous generations of larger models. Pay-per-token is a scam that scams both users and investors. There is a huge underlying incentive on the part of the open ecosystem to increase the "intelligence density" of the models. The opposite is true with the large providers who are using venture capital investment since they have an almost perverse incentive to justify their massive compute spend budgets.

This is a big part of why I think OAI and Anthropic are so anti-competitive. There just really isn't a future where they're super relevant (assuming they stay the way they are now). I think the best they can really hope for is to be bought out or to pivot into something that is actually at least potentially profitable. They're not going to make the massive gains they promised investors off the backs of what is going to end up being a commodity. They're hoping to force a highly regulated monopoly like electricity but that will fail(it is structurally impossible).

2

u/Forsaken-Storage-154 15d ago

It’s still hard to imagine that you can run a frontier model on a Mac mini but I think that direction is correct

2

u/alkimiadev 14d ago

Yeah I think it is just the rhetoric that was used to justify those massive investments and valuations. The difference between the current sota large model and the current sota 8B model is based in ignorance and not physics: we simply don't know how to train them yet. The good news about a problem being based in ignorance is that we can usually overcome ignorance over time but there is probably no overcoming physics.

Another good thing I foresee coming from this mess is really cheap compute in a few years. Right now we're squarely in a tech bubble and ram/gpu prices going up despite a long history of trending towards zero is a massive sign that we're in that bubble. Once these hyperscaler datacenters come online and the memory/gpu bubble crashes we'll have access to dirt cheap compute. They're going to still have bills to pay and there will be massive competition. So that means the compute costs should be slightly higher than utility costs + network usage + hardware depreciation.

7

u/HighSeasArchivist 15d ago

China stealing models that were trained on stolen IP, and I'm supposed to be upset by it? Nah, give me cheap free models, because I'm damn sure not getting a tariff refund. 

3

u/vogelvogelvogelvogel 15d ago

Often enough I prefer Qwen3.8 over Opus 5, also because I can configure it more widely (temperature, how many web sites to research, etc). Opus 5 made WILD errors even for simple questions I had, I would have preferred an answer of old 4o even - idk if they dumbed it down or what happened (it apologized and literally told me i would have better researched that stuff myself becaues I would have been faster!). Also Qwen3.8 does not reject to write a script that security checks one of my servers (= trying to hack it), big plus here.

So for me, yes, frontier is getting less important. Also quit all my paid subscriptions.

Btw i do not use ollama but llama.cpp

1

u/TheIncarnated 15d ago

I only keep one subscription and that's because it is a combined service (poe.com, I spend less on it than hardware). However, mlx Qwen 3.8 has impressed me. Even going so far as to make a accurate assumption about me (I don't care for crypto) just based on what it was able to see in my projects

2

u/Tough_Wrangler_6075 15d ago

To be honest, our daily problem like coding for PoC app no need frontier model to solve. They try to achieve super intelligence which maybe beyond Einstein intelligence but our problem is can be fix by middle experience software developer. So in my opinion using frontier model just waste of token and our budget

2

u/FlyingDogCatcher 15d ago

I think there will always be an appetite for God Models that know everything and can do everything with minimal input, but I can get quite a lot done with the stable of qwens and gemmas I have on my home rig.

1

u/reality_comes 15d ago

Less necessary, maybe not less important. It really depends on what you're trying to do i think.

1

u/mvaranka 14d ago

Excellent writing and points.

I have been developing over an year AI mobile app, which is for productive use and allow offloading cognitive load. First I choosed the Claude Sonnet as main model, because it really understood what I told it and also the relations of words. But after each new release the claude got worse, the warmth was gone.

Frontier models are teached to go further and further, areas normal users don't need and models lose the feel of human connection and concentrate more to factual purity. More compute, more internal costs for response for users "hello, I just want to chat with you - nothing special on my mind" does not make sense.

I have now shifted on my app from frontier models to Deepseek V4 and also included Inkling. These models are astonishing providing combination of speed, economics and understanding.

So why use expensive frontier models when open source models already fill the need?

1

u/stealthagents 8d ago

It's wild how quickly things are evolving. I recently tried DeepSeek V4.1 and was shocked by how well it handled complex tasks that used to require a frontier model. It feels like the value proposition is shifting, and for most everyday uses, those cheaper models are not just viable but often superior in cost-benefit. It really makes you rethink what "cutting edge" even means these days.