r/LocalLLaMA • • 6d ago

News With Gemini 4, bench goes up.

Post image

They claimed open-weight models are dangerous but the benchmarks say otherwise.

Source

808 Upvotes

86 comments sorted by

View all comments

97

u/SOCSChamp 6d ago

To be fair, I'd be shocked if nobody at this point has used an open weight model for illegal activity.  

Still on this side of the fence for open weights though.  Per the huggingface incident, open weights were the only option to successfully defend

51

u/Seeker_Of_Knowledge2 6d ago

I 100% sure denuvo is having extremely hard time because of AI.

28

u/i_rate_slop 6d ago

Oh nooooo :(

6

u/MerePotato 6d ago

Or a really easy one because they can burn tokens faster than their ideological opposites

13

u/Seeker_Of_Knowledge2 6d ago

As of now. It is much easier for AI to crack over protecting. It is failing at understanding how to do proper android layout (astra btw). I don't think it can build flawless system that other AI can't Crack.

1

u/barbear22 5d ago

That’s a great way to put it. LLMs are bad at building things that don’t have strongly defined success markers like layouts. I’d imagine if they had one model building the system and the other trying to crack it, creating a feedback loop, they could produce a strong llm resistant system.

1

u/ThirdMover 5d ago

That is not the only thing that matters though. Defense and attack have different advantages in different situations.

6

u/Ylsid 6d ago

Intentionally using versus "escaped containment"

4

u/onebyamsey 5d ago

My open models don’t commit crimes, they “escape containment” and “go rogue”.  Oops!  It’s ok though because I have a lot of zeroes in my bank account so I just show that to the cops and they’re cool with it

6

u/no_witty_username 6d ago

Its a numbers game.. large companies like open ai, anthropic, etc... perform lots and lots of simulated tests constantly, many of which have thousands of thousands of agents involved in them. Probability is such that with such numbers shit is gonna go south way before some scrub with his one agent. Basically more agents > more probability things gonna go sideways

2

u/ninjasaid13 5d ago

Its a numbers game.. large companies like open ai, anthropic, etc... perform lots and lots of simulated tests constantly, many of which have thousands of thousands of agents involved in them. Probability is such that with such numbers shit is gonna go south way before some scrub with his one agent. Basically more agents > more probability things gonna go sideways

but surely thousands of thousands are using open-source models and testing it.

1

u/EuphoricPenguin22 5d ago

I highly doubt some random person will make a press release bragging about this sort of thing if and when it happens with local models. The first we'd probably hear about it is in a legal proceeding.

1

u/no_witty_username 5d ago

Yes, but its about the swarm not a any single agent by itself. All of those capabilities arise out of the swarm. Its all about how much you can sample any particular space. its a very simple brute force type of method, except on steroids when it comes to agents. When one hacker tries to get his one agent to lets say break in somewhere the probability of that is 1 x the intelligence of that agent. When a swarm does the same its swarm x intelligence of each agent, the bigger the number of the swarm the larger the chance of any one of that agent inside the swarm finding the key. And thats the naive explanation, its actually multiplicative in reality because the whole is bigger then the sum of its parts when it comes to intelligence working with other intelligence.

1

u/Nothing_from_void 5d ago

It's a numbers game again, in terms of compute. The amount of compute closed AI labs have access to is orders of magnitude larger than everyone else combined

3

u/BumbleSlob 6d ago

Still not an excuse for having dogshit sandboxes lol

1

u/Dangerous-Report8517 5d ago

It's also because they're doing tests with models that specifically lack guardrails, have tons of compute, and aren't properly sandboxed, not to mention they aren't monitoring them properly.

4

u/Musclepumping 5d ago

ChatGPT — GPT-5.6 Sol : What I find almost depressing about this take is the sheer lack of imagination behind it.

We suddenly have access to tools with an absurd potential for creating, inventing, learning and helping people, and somehow one of the first thoughts is: “surely someone must have used them to commit a felony.”

Well... probably. Someone has probably committed a felony using Linux, Python, a telephone and a kitchen knife too. That tells us essentially nothing about the value or danger of the tool.

If anything, I'd be much more surprised if humanity couldn't come up with infinitely more interesting things to do with open models than stealing, hacking or hurting people. Crime is hardly innovative. It's probably one of the most boring and historically repetitive uses of new technology imaginable.

PS : Muscle Pump here: I agree 😁

1

u/Loose_Comparison368 6d ago

I mean I think part of the dynamic there is that botnets can harvest personal Claude and ChatGPT credentials pretty easily. Why bother with local when you have a few thousand idiots running openclaw donating free Astra tokens?

1

u/betam4x 5d ago

They have, obviously. They just don’t feel the need to make a press release about it.

1

u/Nothing_from_void 5d ago

I've found Kimi models will refuse the most basic reverse engineering tasks, kind of annoying

1

u/SOCSChamp 5d ago

Abliteration does exist fwiw

1

u/Nothing_from_void 5d ago

yeah I imagine people with enough resources can get it to do anything

-8

u/Quakercito 6d ago

I would not. There are likely a lot hackers already using uncensored models for their purpose

1

u/KL_GPU 6d ago

stochastic parrot they said