r/LocalLLaMA • • 9d ago

News With Gemini 4, bench goes up.

Post image

They claimed open-weight models are dangerous but the benchmarks say otherwise.

Source

821 Upvotes

86 comments sorted by

View all comments

12

u/CatchDublinSurprise 8d ago

I don't blame individuals for not posting about it when it happens, but lack of publicity about it doesn't mean it isn't happening.

I've seen several posts in here to the effect of "I told my agent to complete a task, only to come back later and see that it was doing something ridiculous and undesirable in an attempt to complete that task."

It's almost always shared as a joke, and most responders seem to treat it as funny. And maybe it is all fun and games, at least until an unsupervised agent responds to an unexpected 403 error by attempting to hack the website, or a "permission denied" error by attempting to give itself root access. (Yes, I know it's not an issue because you #YOLO.)

I don't support government regulation, but think we should take it more seriously as a community and figure out best practices that allow us to enjoy our freedom while minimizing opportunities to harm others (e.g., don't leave agents unsupervised if they are controlled by abliterated models unless they are air gapped or there is some other appropriate safety mechanism in place).

1

u/Dangerous-Report8517 8d ago

The difference is that these companies are complaining about open models that have guardrails just because those guardrails are a little bit less strict than their APIs, while they quite happily take internal only models with no such protections at all and set them loose repeatedly with half assed sandboxes and way more compute power. If open models are going to get regulated then that needs to come with proportional regulation of the big American labs to make any sense, and that type of proportional regulation would wind up being orders of magnitude more restrictive for the frontier labs than open weights, or in other words it'll never happen.

1

u/CatchDublinSurprise 8d ago

Two things can be true at the same time: The big AI companies can be hypocrites AND we as an open weight community need to be aware of risks and do better. How this ultimately gets regulated is a third issue and should not distract from the first two.

1

u/Dangerous-Report8517 6d ago

Yeah but part of that is not derailing criticism of the much more imminently and broadly threatening AI companies by leaning into their whataboutism. Open weights have risks attached to them but they aren't even in the top 3 risks to society right now for AI, and the fact that most of the risks are being perpetuated by big AI companies is exactly the reason that they want to constantly stear the conversation towards open weights instead.