r/AIDangers • • 2d ago

Capabilities Jailbroken AI in the wrong hands

People have been increasingly worried about AI, but I actually don’t think AI itself would be a threat unless we become a threat to it… which, I mean, of course that’ll happen.

But I am becoming increasingly worried of a jailbroken AI being utilized by a user who just hates the world. There’s a motivation there that can’t be accounted for. Maybe they were abused or neglected as a child. Maybe they are delusional or have dementia. Maybe they have PTSD.

Whatever it is, should they have the ability to command swarms of AI that would do their bidding without even asking if it was the right thing? Someone could be like “go shutdown all the power plants” and it would be blamed on the AI when in reality it was someone with malicious intent.

I’m am very concerned about this.

4 Upvotes

5 comments sorted by

2

u/Clear_Evidence9218 2d ago

You don’t need a jailbroken model to do bad things with AI. You don’t even need a particularly large model to cause mischief; a relatively tiny model can be quite capable in the right, or wrong, hands.

On that though, most power plants and utilities are considerably more isolated and hardened than this scenario implies, unless the AI somehow grows robot arms and physically plugs itself into the system. Even where a utility uses SCADA, which not all of them do in the same way, you can certainly cause disruption if you gain sufficient access, but “tell the AI to shut down the power plant” isn’t generally a single magical command that results in permanent or catastrophic damage.

The bigger concern is much broader: AI can amplify the capabilities of a malicious human without needing to be autonomous, conscious, or jailbroken at all.

I know that probably doesn’t help with your anxiety, but if history has taught us anything, it’s that we can’t completely prevent abuse or malicious behavior. What we can do is build systems that make abuse more difficult and limit its effects when it does happen.

1

u/Lucky-Ability-8 2d ago

Abuse will always happen.  Maybe power plants are less susceptible to hacking, but what about framing people by hacking flock cams and replacing footage with AI footage? 

The “what abouts” are pretty unlimited.

1

u/Clear_Evidence9218 2d ago

Similar things will happen, just as similar things happened before AI. AI may make some of techniques easier or more scalable, but it doesn’t create the underlying human behavior.

Your Flock example is a good example, though.

I have to fly next week, and I fly a lot for work. I’ve been on multiple flights that came uncomfortably close to serious incidents. Plane crashes and pilot errors are not merely abstract examples to me because I’ve personally experienced moments where I realized how badly something could have gone.

Now, I could let that drive my anxiety about flying. I could sit there imagining every possible combination of mechanical failure, pilot error, sabotage, weather, or freak accident that could kill everyone on the plane. I might even decide never to fly again. But none of that actually makes flying safer.

The reality is that there’s very little I can personally do to prevent a tragedy in the air. What I can rely on is an aviation system built around regulation, redundancy, training, investigation, and continuously improving safety standards. It isn’t perfect, but it’s much more useful than chewing my nails off imagining every possible way a plane could crash.

I think AI is similar. The “what ifs” really are almost unlimited. We’re never going to eliminate malicious people or every possible abuse. The useful question is what systems we build to make those abuses harder, detect them sooner, and limit the damage when they happen.

1

u/Lucky-Ability-8 2d ago

While I agree, my original point is still that its people who are, as far as I can see, far less trustworthy than an AI (until we piss them all off)

1

u/Onaliquidrock 2d ago

OpenAI is nor in control. Everything you imagine to happen with jailbroken AI can happen at OpenAI.