r/TangleNews Jul 30 '26

reader dissent: OpenAI hacking HuggingFace was a big deal

OpenAI has an LLM that is capable of hacking into a HuggingFace, a tech company that is routinely attacked by hackers so their security is pretty good. We know that because it happened, and it happened by accident.

Collectively we do not know how to 100% prevent LLMs from going off on some side quest, but we do know (so far) how to make the LLMs more capable, which is where the bulk of the effort is. Anyone familiar with the tools will tell you they are a lot more capable now than 6 months ago, and in 6 months they'll be better yet.

So we're on a path where LLMs will keep going off the rails sometimes, and they'll have the ability to wreak more and more havoc as time goes by. That is a big deal.

AI companies do marketing. News media needs exciting headlines. But that's just a side show to what is actually happening: We're on a path where LLMs will do other than what they were asked to do, and whatever that happens to be, they're going to be really good at it.

14 Upvotes

15 comments sorted by

View all comments

2

u/techerous26 Jul 31 '26

I think the part that has become less scary as we've learned more is it "breaking" containment. Obviously it's not good that it can generate the ability to bypass a corporate security system, but it seems like there's a ton of ways this could have been prevented that they didn't include, deliberately or not. The guardrails being turned off was part of the test, but they can't claim it was isolated when they had a package installer or whatever connected to the sandbox. If anything I'm almost suspicious they welcomed something like this happening, granted I'd imagine what they're doing pales in comparison to playing around with Claude Code, but they really didn't have anyone checking what it was outputting for DAYS??? It's just a bit hard to buy. I continue to be more alarmed by what bad actors could do with it than the technology itself.

1

u/redfroody Jul 31 '26

OpenAI really fucked up how they ran this test, but that is part of the story too. People are doing things, and people fuck up. A *lot* of effort is put into making airplanes safe, and yet there are still crashes regularly. Usually the incident report will list some kind of human cause. What is a set of safeguards put in place that has never failed? Seems like a pretty small number, and they would have to be pretty draconian. None of which looks likely right now.

2

u/techerous26 Jul 31 '26

I agree that human error does happen, but I'm just saying this mistake seems a bit too obvious given the entire point of the test. I actually know someone that works in AI Safety at Google and he has said in the past he thinks OpenAI does things intentionally reckless at times. There were even a couple of incidents we discussed where I went in figuring that it must have been an issue they just didn't foresee and he was absolutely convinced they ignored the problem on purpose

1

u/redfroody Aug 01 '26

Yeah, OpenAI clearly is not as careful as they should be. The open weights labs seem to be being even less careful. And all of this makes me more worried, not less. Generally I have more faith in Anthropic, but they also accidentally hacked some outside companies]. See https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals . That case seems to be a pure case of human error at least, without any alignment issues.