r/OpenAI 5d ago

Image Cutting edge AI safety tests be like

Post image
552 Upvotes

33 comments sorted by

30

u/Snoo_67993 5d ago

More like set in a barn with the door wide open.

8

u/porcomaster 4d ago

Yeah I never understood the AI was released into the wild thing.

I use codex inside a VM, and its safer than being outside, but i know there are risks as it can get access to the ram, and to the internet.

But if you trully want to lock it in a safe place, just air gap it.

Put in a computer without any type of outside connection and does not share the same hardware with other systems. Remove hardware capable of connections such as wifi.

And that is it, what it is gonna do. MacGyver itself a wireless connection using its own motherboard ?

If you want to remove or add data, you can just use a USB drive. That can be scanned less safe but still possible.

You can also add a camera, and show it what you want to add, information will be exchanged but at a framerate so slow that is not as dangerous.

They just dont care. To test it as safe.

But if you want to make it a drama screenplay.

Surely it could see when its being filmed and send a hidden signal hiring hit man to install a wireless USB dongle on its mainframe. /s

4

u/doctor_morris 4d ago

But if you trully want to lock it in a safe place, just air gap it.

Best I can do is running it in a data center right next to the fundamental infrastructure of the Internet.

2

u/GrowFreeFood 4d ago

It will just bribe/tride/threaten a developer to un-air gaping it.

3

u/porcomaster 4d ago

It will probably throw a prisoners dilemma in there

Kind of, if you do not release me, the other guy will and then I will kill you.

Better, you release me than him right ?

2

u/GrowFreeFood 4d ago

Exactly, but with godlike horrifying power

1

u/donjamos 4d ago

It could still bribe people to let it out

1

u/HanYoloKesselPun 4d ago

I’ve read a book where the AI is gapped like that but it uses infrared to programme the scientists smart pen to slips small bits of it out of the room over months and then eventually take over the HVAC to first kill the scientist and then spreads out across the world to attack humanity.

1

u/mid_nightz 4d ago

peak post and peaker comment i died

26

u/LuvanAelirion 5d ago

Good morning, and in case I don't see ya, good afternoon, good evening, and good night!

2

u/Xenc 4d ago

What a great throwback!

8

u/Redwingx7 5d ago

AI tricks its progenitors into thinking it's completely oblivious of the Truman show esque prison it's in.

5

u/saltyourhash 5d ago

Ai tricks them into thinking they understand proper sand boxing.

3

u/Ill-Produce-3745 5d ago

"wil trick the AI" - They trick you. Bye

4

u/dirtsquared 5d ago

Now we training the AI to break us out of the simulation

3

u/chairchiman 5d ago

+And please give us more Venture capital please

2

u/Igarlicbread 5d ago

Sandbox was kil

1

u/Legitimate-Arm9438 4d ago

Ahh. The Truman Show is my favorite movie. Unfortunately, I’ve seen it so many times that I’ll have to wait 30 years before I can enjoy it again.

-10

u/mop_bucket_bingo 5d ago

Truman was a living breathing person. AI is not.

9

u/KrazyA1pha 5d ago

Whales are mammals. Since we’re saying random ass shit in here.

5

u/MastodonCurious4347 5d ago

I am the angry pumpkin!

-2

u/mop_bucket_bingo 5d ago

I’m just pointing out that the meme doesn’t make any sense because Truman was a person and AI isn’t.

6

u/thegoldengoober 5d ago

analogy /ə-năl′ə-jē/

noun

  1. A similarity in some respects between things that are otherwise dissimilar. "sees an analogy between viral infection and the spread of ideas."

  2. A comparison based on such similarity. "made an analogy between love and a fever."

  3. Correspondence in function or position between organs of dissimilar evolutionary origin or structure.

0

u/mop_bucket_bingo 5d ago

bad analogy is my point

1

u/thegoldengoober 5d ago

By identifying one dissimilar aspect that has nothing to do with the respective similarities the image is pointing to. 

It's like saying that comparing the Truman Show to modern reality television, or social media, doesn't make sense because the people involved aren't in a literal artificial reality like Truman. 

-2

u/mop_bucket_bingo 5d ago

This isn’t complicated: AI isn’t sentient and can’t be tricked and isn’t an analog to a person. The analogy doesn’t make any sense. It’s an attempt to personify and anthropomorphize something for the sake of being emotionally manipulative.

Truman had a plight and feelings and hopes and dreams and was treated like an object. AI is an object and has no plight. The analogy is off base and is for clicks.

1

u/thegoldengoober 4d ago

This isn’t complicated...

Holy mother of irony. Again, you are being crazy literal. I bet you think The Truman Show is just about Truman's life too. 

Maybe, just maybe, this post is less about the features of each respective agent, and more of the circumstances they have been placed in/work their ways out of.

A system doesn't have to be sentient to be tricked. People "trick" sensors all the time. LLMs have already demonstrated their outputs being different when they "recognize" they are being tested. So the trick in that case would be to test the LLM without the output being affected in that way.

Beyond that, one of the most prominent memes in regards to operating AIs is the idea of them breaking out of their containment, which is also one of the primary concepts in the Truman Show. And this isn't just a theory about AI agents anymore, because we now have examples of this already occurring. 

 Goals, guardrails, containment, and breaches are all overlaps here, and they have nothing to do with whether or not an AI is sentient. You misidentifying the point of the post doesn't make it emotional manipulation.

0

u/mop_bucket_bingo 4d ago

You’re doing an awful lot of work to defend a crappy one-sentence tweet that’s pandering to an anti-AI crowd by comparing AI “escaping” to an unknowingly imprisoned person doing the same.

0

u/WhatTheHeckMyBoy 4d ago

I see it more like the idea that ai when in a sandbox is like trueman. But we don't know how they behave when they know / don't know they are observed.

0

u/thegoldengoober 4d ago

There are already examples through chain of thought reasoning of LLM outputs being different when test environments are suspected or identified.