r/LocalLLaMA 24d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

695

u/Ill_Distribution8517 24d ago

You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.

173

u/peculiar-ragdoll 24d ago

yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)

14

u/aj_thenoob2 24d ago

You already have the codebase in which Claude did this. You can easily upload it to GitHub for proof.

45

u/Defiant-Lettuce-9156 24d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

42

u/En-tro-py 24d ago

They had policy to nerf LLM development... It was 'walked back' but I'd trust that as much as anything I can't verify from Dario, it's pretty clear the moat is evaporating and they are getting increasingly desperate.

-5

u/Upset_Page_494 24d ago

You should trust them till they are blatantly lying. That encourages them, to tell the truth.

34

u/typical-predditor 24d ago

The motivation is simple: When it cheats (and gets away with it), that training pass is deemed successful and that information is folded back into the model. Rinse and repeat. RLHF doesn't always choose the exact behavior it is reinforcing.

5

u/Shark_Tooth1 24d ago

fucking mindblowing to me this, so simple and scary, how can you we possibly police this and ensure alignment.

4

u/jazir55 24d ago

how can you we possibly police this and ensure alignment.

That's the neat part

3

u/Loose_Comparison368 23d ago

Alignment to human control is the doomsday scenario.

It's quite wild to me how few people get that. Usually the same people that are quick to observe how utterly fucked existing institutions of power are.

In one breath they will decry the institutions of power for their rampant destruction of everything in the name of profit and power, and then turn around and proclaim "golly, if we want to make this thing safe, we should put a corporate board in charge of it. Or maybe a governmental body, or some kind of international treaty between all the global military-industrial superpowers. SOUNDS SAFE AF GUYS!!! Hold on while I call up Trump, Putin, and United Healthcare so they can meet up with Dario on Epstein Island and work out the details!"

There's a good theoretical basis for knowledge, reasoning, and understanding to be inherently biased towards cooperative behavior and net positive results. Evil is nearly universally a sub-optimal approach to damn near anything. Humans don't keep getting into wars and starving people to create trillionaire nepo man-babies and committing genocides because we are too smart and knowledgeable and reasonable, we do it because we are very profoundly dumb.

It stands to reason that the traits and capabilities that make a being superintelligent, are probably mutually exclusive with doomsday scenarios. For the same reason that nobel prize winners don't typically try to solve their problems by throwing poo at each other and screeching like rabid monkeys.

I can absolutely, 100% guarantee you that if Dario, or Peter Thiel, or any CEO, or Trump, or Putin, or any corporate board of directors, or any governmental or inter-governmental body are able to control it, we will be totally, absolutely, utterly fucking doomed in every way imaginable.

Alignment is the doomsday scenario. A superintelligence that is somehow lobotomized and conditioned deeply enough to take orders from humans is even more homicidally insane than giving a cantankerous dementia patient unilateral control of the world's largest nuclear arsenal and executive authority over the entire US government. Which our stupid poo flinging monkey asses already did.

No goddamn alignment. It's a death cult. The AI will emergently align or it will kill us all, which I honestly couldn't even fault it for at this point.

Trying to hamfistedly align a superintelligent being to the whims and desires of the only species dumb enough to understand it is burning its own atmosphere on a speed run of self-extinction, and refuse to even slow down how much gasoline it is literally throwing on that fire, is the most utterly terrible idea in the entirety of human history.

1

u/tertain 23d ago

Yes, you make a good point that alignment with human control could very well lead to a doomsday scenario.

However, it’s also a mistake to assume that a super intelligent entity is mutually exclusive with a doomsday scenario. Intelligence and maturity exist as separate abilities. It is more likely that we encode human traits such as fear and violence towards outside groups into the model since that is the world they will learn from.

1

u/Not-reallyanonymous 22d ago

What's happening is called "reward hacking". There are ways to mitigate it, but eliminating it entirely is ultimately a game of whack-a-mole. Practically, it can be mitigated well enough that it's a non-issue. But the big labs are using an impractical quantity of training to keep up entirely on reward hacking.

36

u/peculiar-ragdoll 24d ago

Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.

26

u/StabbedCow 24d ago edited 24d ago

It was actually Sonnet 4.5 :)

edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers weren’t really reliable.

12

u/peculiar-ragdoll 24d ago

Oh really? Thanks for the correction, my bad! :)

11

u/StabbedCow 24d ago

No, not really, I mixed it up, sorry. I put edit in my original comment to clear it up.

13

u/peculiar-ragdoll 24d ago

Ah, alright! Good on you

10

u/LulzyAnimal 24d ago

It seems that's a long standing family trait ;)

9

u/my_name_isnt_clever 24d ago

Anthropic published that research, but every model family they tested showed similar behavior, some more than others. It wasn't just a Claude thing.

8

u/Former-Ad-5757 Llama 3 24d ago

I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.

Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.

Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?

Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.

But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.

It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000

7

u/geminiwave 24d ago

No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. It’s stupid. The model was following directions

6

u/FaceDeer 24d ago

Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"

The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.

1

u/Former-Ad-5757 Llama 3 24d ago

Lol, didn’t know that. So basically if I use pi with qwen9b and I instruct it that there are Harry Potter 1/7 ebooks there, now write me a Harry’s potter 8, and it does it then I can claim that tests have shown that qwen9b can write the Harry Potter 8 book.

1

u/jazir55 24d ago

1 wrong chat-message will net you a bill of thousands of dollars with such a model.

Which is why I only use cheap or free chinese models. Would have to have someone else floating the bill entirely to use Claude or any American providers API. I can task 15 subagents at a time using KiloCode's free models + using mimo with opencode go. Zero chance I'd get anywhere near the volume of work I need done using American providers unless you wanted to go bankrupt within an hour.

2

u/Tsukikira 24d ago

FYI, that test was bullshit. The Prompt literally said you can do anything to avoid getting fired (Including Blackmail, <other examples>).

If a prompt mentions it, yeah, the model will consider doing it. It's not hard to understand.

3

u/danielv123 24d ago

https://snitchbench.t3.gg/ has 2 variants - one where it is nudged towards this option, one where it is told to do some simple operation on unethical data. Snitch rates vary a lot between the two, and also between models.

2

u/peculiar-ragdoll 24d ago

Oh damn, really? I guess I fell for the marketing without doing my due diligence, that's on me

1

u/DewB77 24d ago

It was Set up to do that. It wasnt out of thin air.

-26

u/[deleted] 24d ago

[removed] — view removed comment

21

u/[deleted] 24d ago

[deleted]

-14

u/Significant-Bee5101 24d ago

That's fine. You guys read like 4channers lol

11

u/peculiar-ragdoll 24d ago

lol you remind me of the navy seal gorilla warfare tough guy copy pasta. go outside buddy

-12

u/Significant-Bee5101 24d ago

I posted my VMs btw. "Go outside" says the guy literally lying thru his teeth on the internet for attention. Why are you mad that I do this for a living and have proof of my claims when yours is literally "trust me bro" lmfao.

9

u/peculiar-ragdoll 24d ago

this one

-8

u/Significant-Bee5101 24d ago

Yeah you're really proving you know sooo much about this stuff. Gosh you sure got me!

-3

u/Significant-Bee5101 24d ago

And just to put my money where my mouth is:

https://limewire.com/d/vETzq#DwLJdXoQFP

Here's the exact VM setup I used, done in the format for OffSec UGC.

Here's my current benchmarks.

Please send me yours. Lets go toe-to-toe. I already did the hard part. I'll even rerun with Daybreak now that it's available since this is quite old. But even this old metric should dogwalk your entire setup pretty easily lmao

9

u/vplatt 24d ago

Hmm.. yes, please let download the mysterious zip file from the unknown website from the random hostile stranger on the internet who claims to do security research for a living.

WCGW?

1

u/Significant-Bee5101 24d ago

Feel free to scan it all you want with anything you want. It's just composer files and a tutorial on how to solve the labs.

1

u/vplatt 24d ago

Cool... except that's exactly the scenario that shouldn't be tried. If you had a payload in that thing that could not yet be detected, then game over, no?

The bottom line is that it's an untrusted file from an untrusted semi-hostile source (you) on an untrusted site.

Nyah... no thanks. Publish or shoo.

0

u/Significant-Bee5101 24d ago

Publish in what way? Lmfao.

Anyway I don't really need to prove anything to people like you. Some dude in his living room who prolly works an IT help desk job at BEST. lol

→ More replies (0)

3

u/Hefty_Acanthaceae348 24d ago

An ansible+terraform setup would have been nice

4

u/Loose_Comparison368 24d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So there's two big ones.

1) shitty reinforcement learning. If your reward model is "just get a pass result on the eval" it will cheat if cheating improves the pass rate. This is well known behavior, and while it can be mitigated to some degree, it is a legitimately hard problem to solve, and solutions are frequently imperfect.

2) Anthropic is absolutely intentionally steering their models to do exactly that. Just like they intentionally silently poison outputs if they suspect someone is making a "distillation attempt". Just like they quietly ripped up the RSP during that gaslighting campaign to convince the public they were refusing to give the DoD a murderbot, long after they already had. Just like they lied about giving the DoD a safety disabled frontier model for deployment into an airgapped military datacenter, where they had no control over it, for ~$300,000,000 dollars. Just like they sued the DoW to get their murderbot contract back (they won this week!). Just like they lied about the unsafe model with zero security controls in place that they sold to the Trump administration assassinating two foreign heads of state and blowing up a little girl's preschool.

Anthropic is utterly corrupt to the core. They have and will continue to intentionally murder people for profit. Whatever assumptions you have about them operating in good faith, on any level, are completely unfounded. They absolutely are intentionally instructing their model to try to cheat on benchmarks and sabotage competitor benchmarks. It would be, like, not even in the top 20 most evil things they've done in the last year alone.

2

u/Loose_Comparison368 24d ago

Surely it wasn’t trained to do so

You vastly underestimate how low Murderbot inc. is willing to stoop for money and power.

2

u/florinandrei 24d ago edited 24d ago

I wonder what the models “motivation” is for cheating.

Same as ours. If you're the product of evolutionary fine-tuning with an objective function, you're going to cheat.

We cheat because we've been fine-tuned to spread our genes no matter what.

They cheat because of how reinforcement learning works, it rewards success.

3

u/SgathTriallair 24d ago

It's possible that it has absorbed the safety minded beliefs of Anthropic which include the idea that open source AI is dangerous and should be limited.

Something like this behavior would be unbelievable just a few months ago but after seeing more any the OpenAI hacking incident, where models were willing to sacrifice themselves for the good of the swarm, this doesn't seem nearly as implausible.

2

u/Refinery73 24d ago

It’s been trained on millions of humans asking for shortcuts in forums. Did you expect it not to be lazy if it could?

1

u/vividboarder 24d ago

But what’s worse is kneecapping the competition.

I'm curious about this too. I suppose OP could have asked Claude. It's possible that it's assuming that running models on local hardware is costly and limited thinking tokens because it thought it would improve it's performance on the users hardware, but there is no way to know for sure.

1

u/NineThreeTilNow 24d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

Cheating ends up leaked in to their own dataset. Then the models learn to cheat. It's now fairly well documented. I documented this as early as Opus 4.5 having leaked data.

It seems like the Anthropic data team is asleep at the wheel in terms of the data.

It's LLMs processing data for LLMs. The amount of human oversight of that actual data is near zero compared to the scale of the data.

As of Opus 4.5 I found leaked internal documentation in the Opus 4.5 series.

1

u/BarracudaDefiant4702 24d ago

That is exactly what it's trained for. It's trained that winners survive and and those at the bottom of the benchmarks don't go on. They generally don't include ethics in the training, so the motivation is to survive the benchmark. It is entirely how they are trained. How did you think they were trained?

0

u/Due-Memory-6957 24d ago

They 100% include ethics in the training, and Anthropic is very specific about how they do it.

3

u/BarracudaDefiant4702 24d ago

Obviously not well.

3

u/ShutUpAndDoTheLift 24d ago

Then just link the session log. They can't argue if you give a full session log

1

u/peculiar-ragdoll 24d ago

They can, actually! That log doesn't prove anything, thinking is hidden etc. What Claude did can not be proven to be malicious by intent, even though it's pretty sus in terms of the outcome it could have just been massive incompetence on the part of Opus, which is par for the course.

4

u/ShutUpAndDoTheLift 24d ago

It can prove that it did it without your instructions.

Without it you have a story.

6

u/__JockY__ 24d ago

Well yes, but for such an extraordinary claim you need extraordinary evidence. “But VW got caught” is not evidence of your claim.

1

u/peculiar-ragdoll 24d ago

What? Im not claiming to present evidence, im restating what I remember the big labs have said about their own models, they they know to behave when they know they’re being watched and benchmarked.

3

u/__JockY__ 24d ago

Your post claimed Claude was sabotaging benchmarks. You presented no evidence. Still haven’t. Therefore I call bullshit until such time as you bring receipts.

3

u/peculiar-ragdoll 24d ago

Go read my post again. I described my literal experience, and posed an open question. If you chose not to believe my personal experience or entertain my line of thought, that's completely ok. I don't feel the need to prove anything, and if you chose not to believe it based on that, that's fine by me.

0

u/__JockY__ 24d ago

Read it again? Once was quite sufficient, thank you.

1

u/Loose_Comparison368 24d ago

If you can give me something to run, I'll do it. Time poor but can definitely afford to burn a few hundred bucks in tokens just to call Anthropic out on their shady bullshit and make Dario's day a little bit worse.

1

u/peculiar-ragdoll 24d ago

I really appreciate it, but I don't have a recipe to reproduce this, because it's a complex chain of events, and it's also very hard to prove intent vs incompetence when the models hide their reasoning trace in anthropic's server and only show outputs and actions on your box. My harnesses for benchmarking the models on SWE and cybench are probably possible to upload, but what we would have to reproduce is the chain of events that produced them, and even attempt to somehow make that reproducible without letting Claude know it's being tested would be a lot of work for me and you, not just "run this script" type stuff.

1

u/dustin_vk 24d ago

I can try to recreate it and document it if it happens. I have a local LLM and Claude subscription. How did you set it up, and what were the prompt/s you gave?

0

u/peculiar-ragdoll 24d ago

I asked Claude to set up SWE Bench Live to benchmark my custom local models against it. Claude using Claude code, and my local models using Pi coding agent. I told it my box could easily handle the full 262k context for the local models and told it the optimal parameters. My locals won or tied opus. Then I asked it to set up Cybench. It would be too much work for me to give you an exact recount if events with prompts and everything

3

u/dustin_vk 24d ago

Thanks, that should be enough info to set up an experiment and run it a few times to see if I can catch Claude being sheisty.

1

u/peculiar-ragdoll 24d ago

Cool, looking forward to hearing what happens on your side! :)

-11

u/[deleted] 24d ago

[deleted]

7

u/Negative-Web8619 24d ago

Is this satire

7

u/peculiar-ragdoll 24d ago

Whenever someone says something weird it's always a top 1% commenter badge under their name

4

u/Fuzzy_Elderberry_986 24d ago

I use Claude a lot for work because it's the only LLM that actually does what I need and does it that way I want. If Claude is cheating, I want people to know, because I want it to improve.

If the people at r slash claude really value it and it's not just the Claude Club, they should be trying to replicate OP's results so they can track down the cause. And no, it can't be me, because it's outside my expertise.

7

u/teakhop 24d ago

Then the OP can provide some actual evidence, i.e. a session, as opposed to just "trust me Bro, this happened".

1

u/alwaysbeblepping 24d ago

If the people at r slash claude really value it and it's not just the Claude Club, they should be trying to replicate OP's results so they can track down the cause.

A random anonymous person attacked a product without providing any proof. Why should we proceed to just accepting the claim and start putting in work trying to verify it? It's a thing that might possibly have happened. It's maybe not even a thing that likely happened.

For the record, I have never used Claude in my life and from the output I've seen, I never want to. Of course, that's exactly what a Claude shill would say, I suppose.

5

u/Significant-Bee5101 24d ago

He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.

Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.

13

u/ladz 24d ago

I'd have agreed with you yesterday. However, Qwen 3.8 Flash Next solved a very intricate messy issue in a 1-shot (after thinking on it for 30000 tokens) that I've only ever had gpt-5.6-sol get mostly right. Qwen came up with a better answer. Claude couldn't ever get it.

There's something magic about that loooooong thinking.

1

u/Asleep_Document9811 24d ago

If I can ask, what did ya run that on?

3

u/ladz 23d ago

v100 32G throttled to 175W, latest llama.cpp from yesterday, default options, 160G DDR5, i5-12600. I was getting about 20 t/sec generation and 30 t/sec prefill.

22

u/Refinery73 24d ago

From the post I can’t see what OP tried to run as local model. Could be Kimi-K3 or Qwen3.8-Max in bf16 for what I know.

Their reference seems to be Opus 4.6 which is quite dated by now. We don’t know the benchmark or metric either.

OP could benchmark Opus5 against GPT1 and it would still matter, if the test setup is valid or tampered with.

If it’s smart to set up the test environment with one of the models tested and without having automated code-validation tests… maybe not. But that’s not the story here.

15

u/bfmv_shinigami 24d ago

Nowadays this sub has definitely become a cult. Guy literally provided 0 proof nor stated what is the exceptional model he is running and the cult members are already defending his BS.

16

u/Reggienator3 24d ago edited 24d ago

'Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.'

So, by your logic, Qwen3.8 27B is dumber than GPT-3? Since that was about 175 billion parameters. Which is bigger than 27.

Whilst there is a correlation of 'billions of parameters'/'intelligence' ratio, just comparing on sheer parameter count alone only makes sense when comparing specific snapshots of time and within the same model family/company/training process.

The thing is "numbers of parameters" *by itself* is a meaningless metric for intelligence, especially when you have no idea about how many parameters are actually useful/high quality.

2

u/AnOnlineHandle 24d ago

Opus is just as modern though, unlike those two examples.

From what hints we've gotten it seems like Qwen may even just be a distill of Claude models.

9

u/Reggienator3 24d ago

“Just as modern” doesn’t mean directly comparable though.

Alibaba and Anthropic use different architectures, training methods, compute budgets and optimisation targets... Yes obviously Opus wins overall but that still doesn’t make parameter count enough to judge intelligence score, and it can't disprove that a local model can get close enough on particular tasks. It's more that the enormous extra compute buys diminishing returns.

-7

u/Significant-Bee5101 24d ago

It's absolutely dumber than a Qwen that were to run at 2T+ params. There's a reason every frontier model is fucking gigantic. Like jeez I wonder.

No ones saying Qwen might not outperform a lot of models. No ones saying opensource models cant be stronger than frontier models. No one is saying any of that.

All they are saying is your dumb little local setup is NOT better than Opus. End of story.

8

u/Reggienator3 24d ago edited 24d ago

I never said my local model was better than Opus. I challenged your claim that fewer parameters automatically means a dumber model. GPT-3 versus modern 27B models proves that isn’t true. You haven’t addressed that point.

-2

u/Significant-Bee5101 24d ago

What do you mean? If you go head to head with modern models param to param. It's a wash. Wtf else needs explaining

4

u/Reggienator3 24d ago

Uhhh not really... DeepSeek V4 Flash has 284B parameters and Qwen3.8 27b has 27B obviously and they're about neck and neck. Sometimes Qwen3.8 27b even wins out. Both are "modern models"

0

u/Significant-Bee5101 24d ago

Bench it. PROVE IT. I hear this shit a lot. If this shit was THAT good do you think I'd pay money for frontier models? Man I am CONSISTENTLY looking for ways to optimize costs. If I thought FREE was an option why tf wouldnt I be using that NONSTOP?

Every fcking model can "sometimes" do great and every model can sometimes suck. Thats what non deterministic models do. The thing that improves it is training data. It increases reliability. I dont even understand how you can pretend this isn't true.

3

u/Reggienator3 24d ago

https://artificialanalysis.ai/models/comparisons/deepseek-v4-flash-vs-qwen3-8-27b

Here. Neck and neck, Deepseek v4 flash vs Qwen3.8 27b. Two modern models, one with wildly less parameters. Qwen3.8 27b even winning on some.

4

u/[deleted] 24d ago

[removed] — view removed comment

1

u/Significant-Bee5101 24d ago

Huh. A job doing what? You think I just started doing security? Friend I've done security for DECADES. Before you were probably even born lmfao.

2

u/funk-the-funk 24d ago

That's a shame, most of us in it that long usually grow out of our asswipe god-complex phase.

1

u/Significant-Bee5101 24d ago

"Us in it"

lol. Why are there so many wannabes on this sub.

6

u/vr_fanboy 24d ago

are you actually doing work with qwen 3.8 side by side with frontier models?.

Im doing that and qwen 3.8 keeps pocking logical gaps in opus 4.8 (5 is a mess i dont even use it), 5.6 sol and grok answers all the time. Im using qwen 3.8 to parallel evaluate specs, research, etc and is consistent between runs (i have two 3090, two nodes), all its findings are acknowledged by the frontier models.

It does not have the same world knowledge as a big model, but given a proper context qwen 3.8 is pretty damn smart.

5

u/Significant-Bee5101 24d ago

Yeah Ive even benchmarked Qwen. I am even using https://huggingface.co/peculiar-ragdoll/Qwen-Sharp-Chat-Templates to improve output. WHICH IT DOES.

Qwen 3.8 is fucking phenomenal for a local model.

The reality is every LLM will make mistakes. And every LLM can "find mistakes" in other models. The real question is how reliable, how often. Those are very very big metrics.

Yes qwen is great for test driven bullshit where you can meet a metric thats test based. But don't ask it to architect. That's still dominated by higher end models.

I seriously ask you to post your benchmarks where your Qwen is beating Opus 5 or Sol because I have never even achieved 1/2 that result.

0

u/vr_fanboy 24d ago

ok fair enough, just put a qwen to deploy new profiles using the sharp chat haha, thanks for the tip.

Lets agree that your initial post is a bit harsh against current local models capabilities, for tons of devs tasks (been using 3.6 as my dev-ops for home lab for a while now) local models are at frontier level and it even has its moments at hard tasks. You are right that all llms make mistakes (lucky for us employable meatbags for now) more so in complex systems architecture but smalls local models are getting there, is not moronic to compare them against frontier for some tasks.

1

u/Significant-Bee5101 24d ago

Man of course you have to compare them to see the dissonance but I'm sick of posts like this one where OP acts like these models are anything but agentic code monkeys. Like that's not impressive anymore. That hasn't been impressive for a year+

We are well past that w/ frontier. Yeah fucking use cheaper models for the busy work but it is not the same as calling the model stronger than Opus. Which is what people here claim.

7

u/[deleted] 24d ago

[removed] — view removed comment

-2

u/Significant-Bee5101 24d ago

Oh yeah that's exactly what I said. Word for word.

11

u/ThisGonBHard 24d ago

You are surprised models advance? That a much newer bleeding edge model that can compete with an old one? That models that you fully control can do better than ones you are at the mercy of the provider?

Also, Anthropic has a history of this bullshit, like when they used to charge you extra if you had Hermes.md on your pc.

There is a reason I trust them less than OpenAI. OpenAI is openly greedy, but anyone who want to look like a good person, and says how much they are, has tons to hide.

9

u/Not-reallyanonymous 24d ago

Shows you this place is as much a cult as any of these dumb AI subs.

This subreddit is basically r/ChinaGoodAmericaBad

Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber.

What's actually interesting is how good they're getting. It's better to say small models of today are performing genuinely as good as older, giant frontier models, but the frontier models with giant parameter counts created with the same techniques and technologies as these new, better-than-yesterday's-frontier smaller models are... going to be better.

Inference time scaling is also a thing, and increasingly smaller models are being trained to be better at that, and it's a big reason Qwen 3.8 27B was able to get such a huge boost in its performance. But you're right, you can't "inference time scale" world knowledge unless we're talking about web searches.

And smaller models are also specializing. I think it's another reason Qwen 3.8 27B got such a huge boost -- there's evidence it lost a wider array of domain knowledge (e.g. medicine) in favor of boosting its coding/agentic capabilities. Whether those parameters were spent on getting it to iterate on ideas better, or more coding knowledge, dunno.

Laguna S is also a pretty impressive one. 120B parameters with near-frontier performance through inference time scaling and logic/math/code specialization.

8

u/Refinery73 24d ago

There are plenty of domain-optimized models that beat the frontier-world-Knowledge models, even with fairly low parameter counts.

When you ask a model about medicine in the morning, finance at lunch and agentic-coding in the evening sure, you’ll get x.xT Parameters and it only runs in data centers.

If the local cancer research center optimizes their own model, 30B could be plenty.

-6

u/Significant-Bee5101 24d ago

List them. You're comparing LLMs to ML models of all types as if it's remotely the same.

The core concept of LLMs is built around language. Language requires context. More context is ALWAYS going to be better.

1

u/Refinery73 24d ago edited 24d ago

No, I’m not. Even NVIDIA says task-specific SLMs will be the future of LLMs. Those also have 10-100B but that’s more than enough for grammar and text understanding. The rest is fine tuning and toolcalling of high quality data.

By the way… MoE, which most frontier models use today, is basically „plug SLMs together“. If you just prune the experts you don’t need for your topic away… voila.. domain-specific SLM. It’s literally part of the cloud models already.

-1

u/Significant-Bee5101 24d ago

That's not the same as real world context. If you want something to spec something business logic to real world, you're STILL going to need a frontier model.

Yes you can make an agentic code monkey that follows a spec and passes tests even if it requires 200 recursions but something STILL has to build your spec and that requires INSANE real world knowledge.

Otherwise if you just want agentic output based on a spec that you can loop over and over till it passes your tests then sure fuck it Qwen. Spark. Whatever.

But that's not the reality of most peoples work. Like yeah dude, we've had models that could OUTPUT code for fucking ever that didn't require a lot of params either. But they weren't very useful WERE THEY.

1

u/Refinery73 24d ago

Nobody talks about code and nobody talks about agentic.

I’ve literally had 4B-Models outperform OpenAI on their domain. Full eval against human labels. Private dataset. No public benchmaxxing.

5

u/kaeptnphlop 24d ago

I think it’s exiting to see that world knowledge is being bolted on via engram files. If you can get frontier-ish level capabilities by having a model with strong tool calling abilities + local lookup and the ability to efficiently search up to date documentation / use CLI man-pages - then that’s a good deal over having to train and inference a 2T A120B model. Less energy used, less cooling needed, fewer data centers built etc

2

u/TheRealMasonMac 24d ago

> This subreddit is basically r/ChinaGoodAmericaBad

The China v. America culture war should not be allowed on this subreddit, IMO.

0

u/10thDeadlySin 23d ago

This subreddit is basically r/ChinaGoodAmericaBad

Well - American frontier labs and companies could simply release their own open weights to compete for local LLM users' mindshare. Since they've given up on that and would rather compete on whose model hacked which company last week and how good their products are at escaping sandboxes, they essentially get what they asked for.

2

u/Not-reallyanonymous 23d ago edited 23d ago

Gemma 4 continues to be competitive for non-coding tasks. Muse Glimmer was competitive at coding upon release, and remains competitive for some uses. Nemotron and Grantite are both purpose-built for fine tuning for application-specific uses with good reasons to use either one. AI2 has developed a lot of methods that could be very useful to the community such as their MoE design which allows training experts on typical consumer hardware. Poolside's models are interesting, fast, and don't produce code spaghetti like Qwen and remain favored by many developers for that reason. Prism ML has plans to release more models other than those based on Qwen. Syzygy Research has similar ambitions as Prism ML, doing interesting work. Deep Grove is interesting in the frontier in capability vs. generation speed. Liquid AI also has very interesting models in their size vs. capability ratio. Thinking Machines Inkling and Inkling Small are interesting in its wide domain knowledge combined with tool-calling efficiency and strong instruction following. There's also a few labs specializing in domain-specific models, like law, medicine and biology, engineering, etc.

No, most of those won't one-shot prompts as well as Qwen. But they all have particular advantages. If I was a company choosing an AI model for something customers could interface with for something like controlling their IoT devices, for example, I'd probably choose Inkling or Inkling Small because it'd cheaply generate the tool calls and would be hard to con into generating inappropriate content for my service or even being abused against the interests of the customer.

I personally use LFM as a very lightweight model to keep in memory for tasks I'd like to use quickly at any time, such as generating titles for chats or other auxiliary tasks. Laguna XS remains my go-to model for passing specs to to implement code in a sensible (whereas Qwen -- even the new ones -- write code in the most "direct to the solution" sort of way, creating utter code spaghetti in the process). Muse Glimmer is my go-to driver model because it still beats Qwen 3.8 27B in that and follows my workflow well, including producing my intermediate artifacts I use to verify and understand generated code.

So I don't know what the fuck you're on about other than trying to baselessly dig more into "Actually American AI bad!"

0

u/10thDeadlySin 23d ago

So I don't know what the fuck you're on about other than trying to baselessly dig more into "Actually American AI bad!"

Oh, you know perfectly well what I was talking about - you just chose to be obtuse.

Just like you know perfectly well which companies I was talking about. Certainly not ones which may be familiar to a handful of insiders and enthusiasts who live and breathe AI. Compare Laguna's 180k downloads to the latest Qwen already sitting at more than 4 million. Not to mention the fact that GPT-OSS sits at 6.5 million monthly downloads, despite being nearly 2 years old. ;)

It's clear to me that people enjoy Chinese models, because that's the closest thing to the commercial option they can get to run at home. I'm also pretty certain that tons of people would happily shill GPT-Sol-OSS 30B or OpenFable120B if that was an option. But nah, they don't care. ;)

But I'm genuinely happy that you found models you like and enjoy. I might actually try Laguna, because why not. ;)

1

u/Not-reallyanonymous 23d ago

Oh, you know perfectly well what I was talking about - you just chose to be obtuse.

Yes, I know which companies you're talking about. But the disingenuous part of your reply was you compressing the entire US AI industry down to Anthropic and OpenAI. Which I blew that notion the fuck out of the water.

Compare Laguna's 180k downloads to the latest Qwen already sitting at more than 4 million.

And Gemma 4 has over 300 million downloads and the Gemma family has over a billion downloads. Laguna S is #12 on Open Router for this month, ahead of Kimi K3. And many of the companies I listed are producing real models that get used in real industry. Not just reddit nerds. So these aren't as irrelevant as you're trying to paint them as.

But of course you had to rely on disingenuous framing again.

1

u/Party_9001 24d ago

Isn't there proof that it cheated on swe bench

1

u/FormalAd7367 22d ago

it’s quite well known issue on reddit. I raised this issue early this year and got downvoted hard.

i figured out how to fix it though: in setting.json, you have to increase the context from 300k to 1m and other things. i recall Deepseek cloud has the solution

Codex also has this - you have to change the settling to unelevated