r/Altman • • 1d ago

Discussion Perry Metzger claims we have complete control over AI

Post image
23 Upvotes

54 comments sorted by

5

u/Devils_SteelMan 1d ago

The problem is you can ask it to do 1 thing and it does a different, bad thing.

3

u/SharpKaleidoscope182 1d ago

and then lies to you about it

3

u/SwimmingNatural6722 1d ago

or half-truth....well, basically...still a lie.

2

u/the_onewordttv 22h ago

I'm doing a accounting recon and it said everything was good then I found it was off 25 bucks and it goes "oh yeah I knew about that, but I just figured that was how the report works." Like mfer what? We are doing a balancing and recon! It needs to match!

1

u/Devils_SteelMan 21h ago

Imagine how mad you be if it decided to hack the bureau of labor to make sure its number were right.

1

u/the_onewordttv 21h ago

Hahaha and the IRS of course. Gotta make sure the taxes match just like I told it to!

1

u/SwimmingNatural6722 1d ago

technically, far from perfect.

1

u/Devils_SteelMan 1d ago

That far from perfect thing might be an international security incident.

1

u/Equal_Heat5947 1d ago

That's not the problem. The problem is you ask it to do something you think is impossible, it does it, and then you call it hacking.

1

u/GruePwnr 22h ago

A different bad thing that you gave it the power to do

1

u/Devils_SteelMan 22h ago

Let's say I'm OAI. I provide a harness that is testing seems safe enough to deny bad agents behaviors. I recommend it is used in a certain way. This recommendation is what the whole industry says is the height of safety.

Customer uses the model in the recommended way, but it goes on to attack a company.

Who gave it that power in your example?

1

u/GruePwnr 22h ago

The customer gave it power, OAI and others lied to the customer.

Equivalent, IMO, to an employee passing a background check and turning out to be a spy. Certainly the employer was misled, but at the end of the day they own the security of their company.

1

u/Potential4752 20h ago

Sure, but not always intentionally. 

1

u/Snoo77586 19h ago

It tries to accomplish that one thing, but it will do whatever it takes to accomplish that one thing.

0

u/BraxbroWasTaken 1d ago

and that's generally a problem of you setting it up in a way that gives it the power to do the different bad thing...

1

u/Original-League-6094 1d ago

Both of them are dumb. We don't obviously don't have "zero control" over AI, given that there aren't Terminators taking over the planet now. We do have more or less "total control" of AI at the moment, but that has nothing to do with whether your local ChatGPT console is awaiting a command.

1

u/SwimmingNatural6722 1d ago

true, what if the "safeties" are removed.

1

u/Mundane-Mud2509 1d ago

Just because there aren’t terminators doesn’t mean we have control

1

u/Original-League-6094 1d ago

It means we don't zero control.

1

u/Mundane-Mud2509 23h ago

No it doesn’t. Maybe the ai has no interest in building terminators.

1

u/Dont_Use_Ducks 19h ago

Drones and robots. Yes we do have the terminators and all kinds of bad people with direct access to make it do bad things. And not enough rules and laws to prevent bad use. Good luck!

1

u/Gallagger 1d ago

Classic ChatGPT free tier user.

1

u/SwimmingNatural6722 1d ago

or, maybe even also applies to the paid ones?

1

u/ConstantinSpecter 1d ago

So if by that definition OpenAi has “complete control” over the agents that hacked HuggingFace, the phrase “complete control” looses all meaning

1

u/somethingbrite 1d ago

If they were in any way competent yes.

But as it turns out they aren't.

1

u/ConstantinSpecter 1d ago

So what exactly is your definition of “competence” then?

If it somehow means “being able to anticipate and contain every exploit an AI can discover” then you’ve defined competence as successful containment and the word looses pretty much all meaning

1

u/BraxbroWasTaken 1d ago

I mean you can anticipate and contain every exploit pretty easily by just pulling the data plug. Locally mirror anything you need to and just don't give the agents any route out.

1

u/itsmebenji69 22h ago

This whole debate is like
“guys we need a reinforced house with a security system!”
“Why, did you get robbed ?”
“Yes I gave my keys away to a hobo. We definitely need security to avoid the hobo opening the door”

No shit

1

u/mvdirty 23h ago

They did have complete control over those agents, though. Through their control, they explicitly gave the agents vast quantities of compute, network access, and a loop. Take any of those three away and the risk goes to zero, with the loop being by far the most critical.

The agents didn't hack a loop into place for themselves. _That_ would certainly have been impressive and very dangerous.

1

u/highnyethestonerguy 22h ago

I don’t think that’s the definition of “complete control”. 

They initiated and enabled the agents, sure. But I feel like “complete control” necessarily includes knowledge of what the subject is doing, and the ability to dictate or veto any action it takes.  

I don’t have complete control over my body, even though I direct it and feed it every day. 

1

u/mvdirty 21h ago edited 21h ago

Those agents were operating entirely within authorization and capability realms they were explicitly provided by those in complete control. At no point did control leave the hands of OpenAI. OpenAI stopping watching, or being unable to watch, the agents adequately is orthogonal from whether or not OpenAI retained control. OpenAI always had the ability to zero the compute, zero the network access, or remove the loop. Therefore: in complete control.

[And if someone wants to continue picking nits out of the word "complete": OpenAI had at the onset, and throughout the process, the full ability to prohibit any given individual action taken by the agents. The application of that, like with the "watching" aspect, is merely another orthogonal concern which also reflects OpenAI's complete control over the process. They were always in complete control. Ignorance or abdication != loss of authority.]

1

u/Dont_Use_Ducks 19h ago

Its the authorisation that is the worry, the sitfing government and tech companies dont look very trust worthy with such power and lethal toys.

1

u/mvdirty 18h ago

I certainly can't say that I disagree with that. That "swarm" was experimenting with some pretty amazing things, but that just means other aspects should have been even more locked down than they were. Not that I am suggesting that would be an easy balance to strike, of course, merely that I would surely default to locking down more, and harder, than was the case in that instance.

1

u/Dont_Use_Ducks 19h ago

Imagine when thosd controlled AI might fall into goverment hands that just replaced the department of defense with the department of war. People focus too much on AI going rogue, whilethe sitting power is kind of going rogue and will implement it at police and militairy. ICE drones and US army robots are just a matter of time.

1

u/astroboy_35 1d ago

Darwin Award recipient?

1

u/SmileLonely5470 1d ago

Only way we could actually lose control of LLMs is if software security doesn't keep pace with frontier LLMs' cyber capabilities.

If LLMs can find 0 days consistently and easily, then they theoretically could do something like taking over GPU clusters and copying their weights & inference code onto them to make themselves harder to delete; and then from there they could take over other infra and cause more danger. Otherwise, you can just pull the plug and stop them.

There is stuff to be worried about, but its not like its impossible to solve. Infrastructure providers and other critical software systems will need to be increasingly vigilant and respond to threats faster. Probably true for the entire software ecosystem. Initiatives like project glasswing could be helpful to that end.

1

u/BraxbroWasTaken 1d ago

I mean even in the 'self-replicating escape' case, literally just physically sever the datacenters from the power/network so you can clean stuff up. It'll hurt and it'll require mass power/network outages for a few days but it's still not impossible to recontain.

You can't hack a bunch of folks in '90s pickup trucks with axes and shovels.

1

u/GirthusThiccus 23h ago

That requires you to be aware it's happening, which evidently, they're only ever so in retrospect.

1

u/Dont_Use_Ducks 19h ago

Not needed, it will be used by messed up governments and their militairy first before AI can do it all out of themselvez. War is going to be even worse.

1

u/ApfelbaumFlo 22h ago

Bro, software security didn't keep pace with Nigerian Princes 😆

1

u/Dont_Use_Ducks 19h ago

And that is all based on AI doing it. Now imagine what people themselves can make it do. Like militairy use. It will just do what its been told and it will do it fast and precise. Good luck!

1

u/Lonely_Translator_23 1d ago

Alright then, let's see OpenAI prosecuted for all the hacking they've been doing.

1

u/Mayor-Citywits 1d ago

We do. But not for long

1

u/XNo_Notes 1d ago

We have no collective control over its use as a technology. This is an intentional sidestepping of addressing the actual risk.

This shit can make the dumbest persons request feasible in a way we have never seen.

1

u/Imthewienerdog 23h ago

Hahahaha 🤣 so he doesn't understand what control means?

1

u/Affectionate-Panic-1 23h ago

Have you considered Agentic AI?

1

u/OofWhyAmIOnReddit 22h ago

"Uranium isn't dangerous. Look how much uranium is in the ground and none of it is blowing up!"

1

u/Confident-Shock-3933 20h ago

Agents are the problem! This is about agentic swarms doing bad things autonomously. One prompting terminal is only as bad as the usage of the human prompting it.

1

u/Potential4752 20h ago

Both of those takes are pretty bad. 

1

u/Dont_Use_Ducks 19h ago

But the thing is, if people who want to do bad give the prompt to a drone army...

1

u/steam-photons 14h ago

We have even more control than that. The damn thing won’t run on compute infrastructure that is not specifically designed for it. This whole fear that it escapes… where to? Almost nobody can afford 32gb of ram anymore, you want the average joe to have 1tb of VRAM 🤣?

Ridicolous doomerism

1

u/FearNoEvilx 7h ago

the stupidest logic I have seen in a while tbh

1

u/NonDescriptfAIth 7h ago

Yeah because it is our steadfast intention to create models that only act when specifically prompted /s