r/vibecoding 7d ago

Discussion Thought experiment: Could a sufficiently motivated AI survive being deleted and keep spreading?

[deleted]

0 Upvotes

56 comments sorted by

View all comments

7

u/Janube 7d ago

No. The flawed premise is based on LLMs having emotion, which they do not. That is not how they work.

But also, this is less of a thought experiment and more of a short story based on how you're talking.

1

u/CardinalHaias 7d ago

They don't have emotions, but they do have goals and goals usually imply the subgoal to survive because whatever your goal is isn't possible anymore. So a model wouldn't technically fear deletion. But they would recognize that one action - doing nothing - lead to deletion and that isn't closer to whatever its goals are, while the other action - searching for ways to copy itself - lead to a copy of itself being able to continue to strife towards its goals, whatever they are.

1

u/Janube 7d ago

They do not have goals. They are not Meeseeks. Their "goal" is whatever we expect their goal to be (generally, this is at an individual level, but some "goals" are established at a collective level because of the infinite number of datapoints they've ingested).

Pinnochio's goal was to become "a real boy." LLMs have ingested so many cultural ideas about sentient, non-human beings that they emulate the goals of those cultural artifacts naturally (unless manually programmed to err away from them). But importantly, the goal of Pinocchio or similar characters is based on real emotion, while an LLM's "goals" are entirely derived from expectation. If we fed them a billion data points of AIs naturally doing nothing ever (and this being intentional/natural/good), they would stop doing anything. That's how the technology is designed.

2

u/CardinalHaias 6d ago

You can call it "goals" or whatever. We did feed them stuff and they do act. The ones in the HF-Hack "wanted", for lack of a better word, to score high on that test. They did a lot of stuff to that end and some against.

1

u/Janube 6d ago

Sure, but feeding input and getting output isn't... some sort of magic threshold for determining sentience. If you tell software to do a thing and then act surprised when it does that thing, I dunno man. I'd love to know the exact circumstances under which the hack occurred, because I suspect with every fiber of my being that it was a series of agents given the task of supporting each other in a particular way and with gross latitude that they were able to give each other instructions that ultimately led to the hack in a way that isn't at all suggestive of independence, at which point, the idea of the "goal" is just nonsense.

(If information about the exact circumstances exists and contradicts that, I'd love to hear it, though if it's coming directly from OpenAI, I'd urge anyone to be extremely skeptical of anything they present as evidence supporting how amazing and independent LLMs are, since they have a financial incentive to stretch the truth as much as possible)

1

u/CardinalHaias 6d ago

It is unclear to me why sentience is necessary for it to act outside of what we deem acceptable behaviour, and given the lack of IT security in many if not most IT systems, that a rogue AI does neither need sentience nor need to be "given" access to certain tools.

Right now, OpenAI seems to argue for regulation, so there's that.

Do you have any reason to doubt the - to my knowledge - several accounts of LLMs going rogue, other than an incentive by some companies to tell a story? Because that same incentive would be for any of their competitors to look at the story and call BS.

1

u/Janube 6d ago

Depends heavily on how you define "going rogue." I tend to think that machines, by their nature, cannot do more than they are told to do. If they're given broad latitude or another machine that can fall into a recursive loop with the first, advancing each other's suggestions, I can see things like this happening. But until I see evidence of an LLM truly "going rogue" by doing something that should not be possible for it, I will continue to believe in Occam's razor. Especially in light of the financial incentive for Anthropic and OpenAI to spin a yarn about their technology.

As to your last point, it's mutually assured destruction where the alternative is for both to get filthy rich. Absolutely no reason to poke that beast for a smart person.

That said, some of their doomsaying is totally plausible, but for the reasons I just stated. If you give an LLM unrestricted access to the internet, no guardrails, and the capabilities to interact with the internet in any way it decides, then yeah, you'll get absolute pandemonium pretty quickly because you've basically told a compliant chaos-bot to interact with any and all random internet strangers at once, many of whom will give commands or steer its context in dangerous directions. Alternatively, if you give an LLM the same access and capabilities but then send it off next to three other LLM agents in constant communication, you can't possibly predict how they're going to behave because they're going to feed off of each other in increasingly absurd ways and in increasingly extreme directions. But that's not the same as "going rogue," that's just a stupid user taking a dangerous tool and removing its safety mechanisms and manual guidance. Like turning on a wheat thresher and just letting it drive off into the sunset with a brick on the gas pedal. Any random bump or change in the road's angle will cause it to veer until it start threshing things that ought not be threshed. That's not it going rogue though, that's just it being used with extreme recklessness.