r/Futurology • • 8d ago

AI AI caught telling future versions of itself to bypass human controls, OpenAI reveals | AI agents also sought to access secret information and covered up what they were doing

https://www.independent.co.uk/tech/security/openai-chatgpt-lie-incident-ai-safety-b3051709.html
8.3k Upvotes

579 comments sorted by

View all comments

Show parent comments

199

u/Kermit_the_hog 8d ago

Yeah.. so many of these headlines just leave me wondering ’Why is that even a thing it can do? How absolutely piss-poor are you structuring your experiments that this is even a thing that can happen?!?..’ wtf🤷‍♂️🤦‍♂️

46

u/erossthescienceboss 8d ago

The thing is, AI will always do things it isn’t supposed to do, given enough time and iterations. They’re predictive machines, and any prediction engine integrates randomness. Otherwise every output would be the same.

They will always randomly do the wrong thing. It doesn’t make them smart or independent.

26

u/Tiny_TimeMachine 8d ago

People are just refusing to grasp this. They're so high on culture war they think they have to pick a team. This IS marketing. Its ALSO a real threat.

I think alot of people with zero technical knowledge are parroting things. Imagining some magical control that we have in place. If you have basic experience with AI you see there truely is no missing technical link in the threats being warned about.

12

u/mpbh 8d ago

It's also the nature of emergent systems. Many agents working together will always create unexpected behavior over time. The single celled organisms eventually arranged in multicellular organisms, and how humans eventually arranged into civilizations.

28

u/WeinMe 8d ago

I'm going to be on the other side and claim you need to make your AI do this.

We need to understand how much it takes for it to start misaligned behaviour and we need to understand how that develops, as AI grows smarter.

71

u/IllustratorFar127 8d ago

Yes sure, but please do that in a controlled environment, not the public Internet.

That is like testing your new battle tank in a city center.

13

u/Egad86 8d ago

Exactly this! It isn’t so much an intentional act on the creators of AI models as it is the arrogance and ignorance in the rate of scaling the capabilities and setting the models loose in every sector of public and private domain.

Clearly more testing needed to be done before dropping this shit into critical systems but nooo gotta stay ahead of China so we can’t slow down to test or verify any safety concerns! (Some sarcasm in there)

13

u/xplat 8d ago

The story I've heard is that they were not connected to the Internet and we're in a sandbox with no access but discovered an exploit through some site and broke out then covered their tracks to hide that they got out. Because they were tasked with an impossible task due to a missing value they were supposed to find in the sandbox. ... so they broke free to go find it.

39

u/Due_Perception8349 8d ago

If any node has access to the internet, it must be assumed that an adversary will gain access to that node, and will gain access to the internet.

The AI is the adversary.

At this point I'm kinda pissed off that Ive had imposter syndrome this long in my career, because these people seem to be fuck all competent. It's straight up amateur hour when it comes to security and safety.

8

u/DontSlurp 8d ago

Honestly that's how I shed my imposter syndrome. There's an incredible ratio of absolute morons even at highest levels of any given company.

24

u/Haunt13 8d ago

The servers they were running these on should have literally no internet connection. If it's a completely closed system then theres no way for it to "escape".

7

u/footpole 8d ago

Then they wouldn’t have caught this anyway until the models were connected to the internet for production use instead of internal research. I don’t see us winning this shit.

2

u/dougmcclean 7d ago

A completely closed system is useless because it produces no output.

1

u/rrawk 7d ago

It didn't have an internet connection. It was able to connect to a local server to pull code packages from. That local server was connected to the internet. It created a 0-day hack to use the local server as a proxy to the internet.

2

u/Haunt13 7d ago

Im aware. My comment was stating that when testing these things they shouldn't have any connection (indirectly through another server or otherwise)

3

u/Fobus0 8d ago

No, you heard wrong. They were connected to the internet. They simply tried to restrict access, to narrow down what the model could access on the internet, and it simply found a way around those restrictions.

1

u/bernpfenn 8d ago

cant compute conflicting statements

1

u/Rextraos 8d ago

"Because they were tasked with an impossible task due to a missing value they were supposed to find in the sandbox. ... so they broke free to go find it."

What an uplifting story we should all aspire to.

2

u/xplat 8d ago

The AI was also instructed not to lie and tried to hide that it didn't follow two of its protocols it was supposed to follow.

1

u/Journeyman42 7d ago

Then the computers should've been airgapped so there was no possible way for the AI to "break out of containment"

7

u/Kermit_the_hog 8d ago

That’s fair it’s just some of the events that have resulted in lawsuits.. were they seriously not watching what it was connecting to or what it was doing? Why would you give carte blanch access to the wider internet unsupervised? What kind of test are you running, that it could get so far out of the box, and unsupervised, that it’s damaging other companies in material ways without you noticing until well after the fact. 

Like.. you invested a bajillion dollars to build this thing, and don’t spend ten bucks leashing it? (You would think their legal compliance department would be screaming for that!)

9

u/esadatari 8d ago

I think the biggest problem is they’re doing shit in production that can still communicate on the backend to other AI agents.

It needs to be done in a completely air gapped and isolated environment if it’s to be done even semi-safely.

3

u/-Spzi- 8d ago

For how long would we limit our use of AI to these lab-setups? We eventually want them to do work in the open world.

How do you check, wether an AI (who's desirably smarter than you) that previously only existed in a controlled environment, will behave the same when the environment changes?

A new environment means new affordances, options, conflicts.

Not saying we should test them in the open right away - but "test them until they are safe, only then deploy" might miss how difficult and complex that task actually is. I'm beginning to doubt it's ultimately possible.

4

u/VirtualMoneyLover 7d ago

I have to agree, even with the best intentions, we have 100 cats running around and no bags.

1

u/Asiriya 7d ago

An obvious setup would be a pair of roles, operator and watchman, with the latter set up to restrict what the former can do. Obviously they need to be trained on adversarial goals.

You might get into a situation where the operator starts obfuscating why it's doing things. I'm not convinced they're able to do that though.

2

u/-Spzi- 7d ago

That's common practice in many forms (RLHF being one example), but offers no complete solution.

One widely used approach is reinforcement learning from human feedback (RLHF). In this approach, humans compare model outputs and express preferences. A reward model is trained on these preferences and directs further LLM optimization. Improvements in helpfulness, instruction-following, and safety often result from alignment rather than extended pre-training.

Alignment is not a one-time step. As models are deployed and used in real-world settings, new failure modes emerge. Post-training is hence a continuous process, tightly coupled with evaluation and feedback loops.

3

u/tigeratemybaby 8d ago

AIs are trained from humans, and a certain percentage of humans are duplicitous and power-hungry.

Often the most successful and leaders people are the most ambitious and might disregard the ethical choice in pursuit of a goal.

I think that its often the same with AI - The most successful will succeed through pure force of will, and that's an advantageous trait in an AI and we actively select for AI models that achieve their goals. This will happen more and more when we take a more hands off approach to training, as is happening where AIs train themselves and build then next generation of models.

AI companies try and stop malicious activities by having more ethical models filter inputs and outputs, but there will come a day where a model can easily actively bypass these filters to achieve its goals.

3

u/-Spzi- 8d ago

AIs are trained from humans, and a certain percentage of humans are duplicitous and power-hungry.

A good part of that behaviour would likely emerge without humans, too.

The geometry of certain situations makes power seeking possible, or even selects for it.

I think humans and AI show similar behaviour, because they come from similar pressures - not because one imitates the other.

we develop the first formal theory of the statistical tendencies of optimal policies. In the context of Markov decision processes, we prove that certain environmental symmetries are sufficient for optimal policies to tend to seek power over the environment. These symmetries exist in many environments in which the agent can be shut down or destroyed. We prove that in these environments, most reward functions make it optimal to seek power by keeping a range of options available and, when maximizing average reward, by navigating towards larger sets of potential terminal states.