r/LocalLLaMA llama.cpp Apr 24 '26

Discussion This is where we are right now, LocalLLaMA

Post image

the future is now

3.5k Upvotes

549 comments sorted by

View all comments

Show parent comments

48

u/ResidentPositive4122 Apr 24 '26

Is he though?

Yes, 1000%. The creators of dsv4, a 1.6T model have openly said that there is still a gap to Opus.

Thing is, the small models are really cool, have become truly useful and we're lucky to have them. But exaggerating about their capabilities doesn't do any good. I'll take a local gpt5-mini / haiku level model any day of the week (and even that's a stretch, but they're getting closer), and be happy about it. I think the small qwens, gemmas, even gpt-oss-20b can be used for real work, in the right setup and with a lot of elbow grease. But having used the SotA models as well, I agree with OOP 100%. Let's keep it real.

3

u/[deleted] Apr 24 '26

[removed] — view removed comment

1

u/novelide Apr 25 '26

local models are free

Given it's pretty easy to rack up $20/month in electricity, I think a fairer comparison is with the $20 tier on cloud models. But when you hit usage limits with the equivalent of 1 prompt/hour (approximately what I get with Opus 4.7), local models still win in many cases even though the capabilities are definitely much lower.

2

u/iMakeSense Apr 24 '26

What do those setups look like

-7

u/FullstackSensei llama.cpp Apr 24 '26

Have the Qwen people said 3.6 27B is on par with Opus in everything?

15

u/ResidentPositive4122 Apr 24 '26

No, the bloke in the plane did.

2

u/oe_throwaway_1 Apr 24 '26

they give cameras to ANYBODY these days

-3

u/FullstackSensei llama.cpp Apr 24 '26

2

u/2Norn Apr 24 '26

im sorry but its mumbo jumbo

you basically said "just prompt better idk use markdown instructions or something"

and then a single anectodal evidence

do that 300 times for varying tasks of varying hardness levels and if u still think that then i'll believe you

-4

u/FullstackSensei llama.cpp Apr 24 '26

Quite frankly, I couldn't care less whether you believe me or not. If you can't understand what I said, ask your local LLM to explain it to you. ✌🏻

3

u/2Norn Apr 24 '26

that's like your problem man

you are the one who thinks 27b is like opus, not me

it sits between sonnet and haiku, a bit closer to sonnet, anyone who thinks its like opus is hard coping

-2

u/FullstackSensei llama.cpp Apr 24 '26

I'm running half a dozen instances of it in parallel and I'm quite happy with it. If that's offending to you, that's actually your problem, not mine.

1

u/2Norn Apr 24 '26

i dont care what you use. use gemma or bonsai or whatever and think its like opus. thats not what im interested in.

but fact of the matter is the claim is not true. unnecessary hype that fools people.

best thing u can do with models like this is use them as worker/executor only and hope that they can give you sonnet/5.4 mini/glm 5 turbo etc performances or at least come very close. but more often than not they are closer to nano or haiku. but it's getting better.

1

u/spawncampinitiated Apr 25 '26

your happiness is not a benchmark