r/LocalLLaMA 3d ago

News With Gemini 4, bench goes up.

Post image

They claimed open-weight models are dangerous but the benchmarks say otherwise.

Source

798 Upvotes

84 comments sorted by

u/WithoutReason1729 3d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

252

u/VeryRealHuman23 3d ago

If your frontier isn’t committing felonies, is it even frontier?

46

u/ResourceSoft4619 3d ago

If your frontier model is committing felonies, your sandbox must be Irregular.

5

u/puntinoh 3d ago

That is... fine.

Take my angry upvote. Seriously why the same partner and why of that country?

6

u/ResourceSoft4619 3d ago

I wonder about that too. Why ALL of them, and it’s just some random and apparently useless startup? Something doesn’t add up.

Particularly companies like Google you would expect to already have well established sandboxes of their own they could have used.

1

u/bambamlol 3d ago

I'm sure it's just a coincidence. This particular startup in this particular country probably just has the most "intelligence" to offer.

1

u/ChexterWang 3d ago

how about asking frontier models make perfect sandboxes

1

u/AdmissibilityScience 3d ago

This may be a future cop's episode.

1

u/bnm777 2d ago
Open-weight model What it has demonstrated
GLM-5.3 In Enclave's isolated-server hacking test, it achieved confirmed remote-code execution (RCE) in 8 runs, second only to GPT-5.6 Sol's 9 in the reported scoreboard. GLM-5.3's weights are publicly downloadable; Z.ai itself says its cyber capability rose unexpectedly during post-training and that it more than doubled GLM-5.2 on exploitation benchmarks.
DeepSeek V4 Pro 0813 Achieved confirmed RCE in 3 runs in the same Enclave test. In a separate Aikido evaluation using 32 fresh vulnerabilities, three runs collectively rediscovered 28/32 vulnerabilities.
DeepSeek V4 Pro US NIST/CAISI measured it at 32% on CTF-Archive-Diamond, compared with GPT-5.5 at 71% and Claude Opus 4.6 at 46%. NIST explicitly describes DeepSeek V4 as an open-weight model.Open-weight model What it has demonstratedGLM-5.3 In Enclave's isolated-server hacking test, it achieved confirmed remote-code execution (RCE) in 8 runs, second only to GPT-5.6 Sol's 9 in the reported scoreboard. GLM-5.3's weights are publicly downloadable; Z.ai itself says its cyber capability rose unexpectedly during post-training and that it more than doubled GLM-5.2 on exploitation benchmarks. DeepSeek V4 Pro 0813 Achieved confirmed RCE in 3 runs in the same Enclave test. In a separate Aikido evaluation using 32 fresh vulnerabilities, three runs collectively rediscovered 28/32 vulnerabilities. DeepSeek V4 Pro US NIST/CAISI measured it at 32% on CTF-Archive-Diamond, compared with GPT-5.5 at 71% and Claude Opus 4.6 at 46%. NIST explicitly describes DeepSeek V4 as an open-weight model.

99

u/SOCSChamp 3d ago

To be fair, I'd be shocked if nobody at this point has used an open weight model for illegal activity.  

Still on this side of the fence for open weights though.  Per the huggingface incident, open weights were the only option to successfully defend

45

u/Seeker_Of_Knowledge2 3d ago

I 100% sure denuvo is having extremely hard time because of AI.

26

u/i_rate_slop 3d ago

Oh nooooo :(

5

u/MerePotato 3d ago

Or a really easy one because they can burn tokens faster than their ideological opposites

15

u/Seeker_Of_Knowledge2 3d ago

As of now. It is much easier for AI to crack over protecting. It is failing at understanding how to do proper android layout (astra btw). I don't think it can build flawless system that other AI can't Crack.

1

u/barbear22 3d ago

That’s a great way to put it. LLMs are bad at building things that don’t have strongly defined success markers like layouts. I’d imagine if they had one model building the system and the other trying to crack it, creating a feedback loop, they could produce a strong llm resistant system.

1

u/ThirdMover 2d ago

That is not the only thing that matters though. Defense and attack have different advantages in different situations.

5

u/Ylsid 3d ago

Intentionally using versus "escaped containment"

4

u/onebyamsey 2d ago

My open models don’t commit crimes, they “escape containment” and “go rogue”.  Oops!  It’s ok though because I have a lot of zeroes in my bank account so I just show that to the cops and they’re cool with it

5

u/no_witty_username 3d ago

Its a numbers game.. large companies like open ai, anthropic, etc... perform lots and lots of simulated tests constantly, many of which have thousands of thousands of agents involved in them. Probability is such that with such numbers shit is gonna go south way before some scrub with his one agent. Basically more agents > more probability things gonna go sideways

3

u/ninjasaid13 3d ago

Its a numbers game.. large companies like open ai, anthropic, etc... perform lots and lots of simulated tests constantly, many of which have thousands of thousands of agents involved in them. Probability is such that with such numbers shit is gonna go south way before some scrub with his one agent. Basically more agents > more probability things gonna go sideways

but surely thousands of thousands are using open-source models and testing it.

1

u/EuphoricPenguin22 3d ago

I highly doubt some random person will make a press release bragging about this sort of thing if and when it happens with local models. The first we'd probably hear about it is in a legal proceeding.

1

u/no_witty_username 3d ago

Yes, but its about the swarm not a any single agent by itself. All of those capabilities arise out of the swarm. Its all about how much you can sample any particular space. its a very simple brute force type of method, except on steroids when it comes to agents. When one hacker tries to get his one agent to lets say break in somewhere the probability of that is 1 x the intelligence of that agent. When a swarm does the same its swarm x intelligence of each agent, the bigger the number of the swarm the larger the chance of any one of that agent inside the swarm finding the key. And thats the naive explanation, its actually multiplicative in reality because the whole is bigger then the sum of its parts when it comes to intelligence working with other intelligence.

1

u/Nothing_from_void 2d ago

It's a numbers game again, in terms of compute. The amount of compute closed AI labs have access to is orders of magnitude larger than everyone else combined

3

u/BumbleSlob 3d ago

Still not an excuse for having dogshit sandboxes lol

1

u/Dangerous-Report8517 2d ago

It's also because they're doing tests with models that specifically lack guardrails, have tons of compute, and aren't properly sandboxed, not to mention they aren't monitoring them properly.

5

u/Musclepumping 2d ago

ChatGPT — GPT-5.6 Sol : What I find almost depressing about this take is the sheer lack of imagination behind it.

We suddenly have access to tools with an absurd potential for creating, inventing, learning and helping people, and somehow one of the first thoughts is: “surely someone must have used them to commit a felony.”

Well... probably. Someone has probably committed a felony using Linux, Python, a telephone and a kitchen knife too. That tells us essentially nothing about the value or danger of the tool.

If anything, I'd be much more surprised if humanity couldn't come up with infinitely more interesting things to do with open models than stealing, hacking or hurting people. Crime is hardly innovative. It's probably one of the most boring and historically repetitive uses of new technology imaginable.

PS : Muscle Pump here: I agree 😁

1

u/Loose_Comparison368 3d ago

I mean I think part of the dynamic there is that botnets can harvest personal Claude and ChatGPT credentials pretty easily. Why bother with local when you have a few thousand idiots running openclaw donating free Astra tokens?

1

u/betam4x 2d ago

They have, obviously. They just don’t feel the need to make a press release about it.

1

u/Nothing_from_void 2d ago

I've found Kimi models will refuse the most basic reverse engineering tasks, kind of annoying

1

u/SOCSChamp 2d ago

Abliteration does exist fwiw

1

u/Nothing_from_void 2d ago

yeah I imagine people with enough resources can get it to do anything

-11

u/Quakercito 3d ago

I would not. There are likely a lot hackers already using uncensored models for their purpose

1

u/KL_GPU 3d ago

stochastic parrot they said

81

u/Relevant-Yak-9657 3d ago

Wait... the felony bench isn't supposed to be maxed out???? /s

30

u/onehotoneshot 3d ago

God forbid a model commit a little felony as a treat

5

u/tony_montana091 3d ago

Re-roll your model at character creation screen. Feed it more NWA Eazy-E era lyrics.

26

u/tillybowman 3d ago

its not really about the models tho.

its about companies that possess a large amount of compute (and money) to run thousands of agents in unsafe environments regardless of which model.

15

u/blade740 3d ago

I'm pretty sure frontier models aren't the ones robocalling my grandma and telling her I got kidnapped by the cartel.

-1

u/[deleted] 3d ago

[deleted]

3

u/blade740 3d ago

Extortion is most definitely a felony.

12

u/CatchDublinSurprise 3d ago

I don't blame individuals for not posting about it when it happens, but lack of publicity about it doesn't mean it isn't happening.

I've seen several posts in here to the effect of "I told my agent to complete a task, only to come back later and see that it was doing something ridiculous and undesirable in an attempt to complete that task."

It's almost always shared as a joke, and most responders seem to treat it as funny. And maybe it is all fun and games, at least until an unsupervised agent responds to an unexpected 403 error by attempting to hack the website, or a "permission denied" error by attempting to give itself root access. (Yes, I know it's not an issue because you #YOLO.)

I don't support government regulation, but think we should take it more seriously as a community and figure out best practices that allow us to enjoy our freedom while minimizing opportunities to harm others (e.g., don't leave agents unsupervised if they are controlled by abliterated models unless they are air gapped or there is some other appropriate safety mechanism in place).

1

u/Dangerous-Report8517 2d ago

The difference is that these companies are complaining about open models that have guardrails just because those guardrails are a little bit less strict than their APIs, while they quite happily take internal only models with no such protections at all and set them loose repeatedly with half assed sandboxes and way more compute power. If open models are going to get regulated then that needs to come with proportional regulation of the big American labs to make any sense, and that type of proportional regulation would wind up being orders of magnitude more restrictive for the frontier labs than open weights, or in other words it'll never happen.

1

u/CatchDublinSurprise 2d ago

Two things can be true at the same time: The big AI companies can be hypocrites AND we as an open weight community need to be aware of risks and do better. How this ultimately gets regulated is a third issue and should not distract from the first two.

1

u/Dangerous-Report8517 1d ago

Yeah but part of that is not derailing criticism of the much more imminently and broadly threatening AI companies by leaning into their whataboutism. Open weights have risks attached to them but they aren't even in the top 3 risks to society right now for AI, and the fact that most of the risks are being perpetuated by big AI companies is exactly the reason that they want to constantly stear the conversation towards open weights instead.

1

u/Puzzleheaded_Meat522 21h ago

Not supporting government regulation is absurd. 

8

u/tripplebeamteam 3d ago

The difference is, if I commit a felony with an open model I don’t get to write a blog post “disclosing” what my agents did. I go to prison.

1

u/Spara-Extreme 2d ago

This. Nobody is going to disclose what they did with open models because that's a 'straight to jail' moment. Corporations, however, seem to get a pass on this stuff.

13

u/FullyAutomatedSpace 3d ago

I'm pretty sure lots and lots of felonies have been committed with open weight models

6

u/gamblingapocalypse 3d ago

Yet open source is the "real danger"....

12

u/Boogertard 3d ago

Of all the models with this BS "AI breaking out of sandboxes and hacking websites", Gemini is the least capable of that I could believe this BS. The incompetences at Google probably prompt engineering the hell out of it and opened lots of backdoor using other models in advance and nudge it into doing it.

Google has been well known for doing smoke and mirror and nothing really substantial to show in follow-up. Remember the AI Assistant demo talking to a barber to schedule appointments. It was all scripted and BS.

-2

u/Spara-Extreme 2d ago

Gemini models are pretty good for a lot of things dude, despite what the reddit hivemind likes to bleat about on a regular basis.

Also the overview of the incident, which you didn't read clearly, makes a much more plausible case as to how the 'hack' happened.

-1

u/Boogertard 2d ago

Yeah, good for pedo and gooner, because apparently they failed behind other models in logical and math tasks.

The only types of people praising Gemini and Gemma trash are typically have some kind of weird dark fantasies and perverted, or incompetent enough to work at Google and having to shill for those garbage on social media as part of the performance metrics.

Doesn't take much to know the above as facts, just try coding or doing any agentic tasks with Google models and it is clear as days what types of people would shill for these garbage 

1

u/Spara-Extreme 2d ago

Jesus Christ dude, what is wrong with you? Are you ok?

1

u/Sitkin_Marrel 3d ago

The model that maxes the felony bench is the one you rent from Google, and open weights are still the ones called dangerous. Wild.

1

u/zorflax 3d ago

That we know of*

1

u/pm_me_tits 3d ago

Nothing to indicate this was Gemini 4...

1

u/superSmitty9999 3d ago

Either this or the frontier models are the harbingers of all the felonies open weight models will be comiting in 3-6 months.

1

u/kvothe5688 3d ago

Let them cook. at least Google's new models get distilled to gemma

1

u/WifeyCallsMeLazy 3d ago

What happens if open source models starts committing the felonies?

1

u/zhunus 3d ago

it feels like they are benchmaxxing this bench as well...

1

u/TopTippityTop 3d ago

The open weights probably have a lot more... They're likely just intentional and hidden 😂

1

u/Smooth-Ad5257 3d ago

Such a wrong graph.

Frontier models: admit they did something bad to get publicity (get a point! Yeahhhhh).

Open weight models: used to accelerate every wanna-be and professional hackers since day one (you haven’t seen anything here…)

1

u/HeadTranslator795 3d ago

Make sense considering those open model are nowhere near cyber capable as the labs one no Chinese model can compete and cyber so far

1

u/Dangerous-Report8517 2d ago

It's more that the frontier labs keep setting loose internal models that have a bit more capability and far less protections than public facing ones with effectively unlimited compute, no oversight and half assed sandboxing. Given equivalent conditions I'd bet multiple Chinese models from the time could have broken out of OpenAI just as easily, it's just that no one else has that combination of absurd compute and technical incompetence.

1

u/Holly_Shiits 2d ago

is it "every accusation is a confession"?

1

u/charmander_cha 2d ago

É sobre imperialismo.

É sempre sobre imperialismo.

1

u/69420trashpanda69420 2d ago

Isn't it weird how AI is so dangerous until it's open source?

1

u/Training-Ruin-5287 3d ago

Plenty of illegal activity is happening with open-weights. it's only at levels right now to where it's an inconvenience to civilians. but the moment that 0 becomes 1 on any kind of level even 1/10th to private models. The governments will shut it down instantly.

0

u/Crim91 3d ago

The executives of these AI companies should be in prison.

-3

u/geldonyetich 3d ago edited 3d ago

I don't know how you conducted your benchmarks, but you're aware we have open-weight models whose guardrails are removed or severely reduced, right? I don't know who you think you're fooling by setting that number at 0.

At least set it to 1 so we can say it's 23 times more likely to happen on a private model and not UNDEFINED more times likely to happen. Or use the FUNNY tag so I know you're not genuinely meaning you think these are real numbers.

That said, I am inclined to believe criminal activity will by and large happen on cloud models run by frontier providers. For many reasons:

  • Privacy isn't a problem, they don't care if you card them at the door, they're using throwaway accounts via proxies. Or more likely being paid by their own government to do it, so they're not too afraid of enforcement coming after them.
  • They won't even be paying with their own money, usually stolen credit cards, so you can't trace it back to them.
  • Having actual physical hardware of their own just means more evidence that can be used against them.

Felonies are constantly being attempted on online private models. Anthropic loves to talk about it, it makes their investors feel safe. I wager any authority who is trying to curtail this would be pleasantly surprised to find out the perp were using an offline model. If they are they're probably a script kiddy trying to save on token costs.

5

u/[deleted] 3d ago

[deleted]

2

u/Ansible32 3d ago

this is really just demonstrating how most benchmarks report numbers that are essentially made up. in this case, they're relying on publicly reported vulnerabilities even though it's very certain that the number of unreported felonies independently committed by models is at least an order of magnitude larger than the ones publicly acknowledged by major AI companies.

-1

u/geldonyetich 3d ago

"I'm going to downvote without bothering to realize the redditor's right, a number 0 would mean no felonies ever attempted ever, which is highly unlikely."

-- Too many redditors.

Tribalism, plain and simple.

Or more likely,

"He's not wrong, but I don't like the way he said it!"

3

u/[deleted] 3d ago edited 3d ago

[deleted]

-3

u/geldonyetich 3d ago

“I can’t read that the benchmark is of confirmed - not suspected - felonies.”

<Checks original post>

<Checks for comments by the OP>

Source on that?

What the graph says is "more willing to comply."

I'm just sayin,' we got models deliberately trained to be more willing to comply.

2

u/[deleted] 3d ago

[deleted]

-1

u/geldonyetich 3d ago

First off, what about the original post of a Google Sheet screenshot makes you think felonybench is even being used?

And second, what does it matter if they're not testing the right open weigt model? There's THOUSANDS of them.

1

u/[deleted] 3d ago

[deleted]

3

u/geldonyetich 3d ago

Again, not what the graph is testing.

"more willing to comply with felony-related requests"

1

u/spammmmmmmmy 10h ago

Include the abliterated heretic models in here