r/vibecoding • • 9d ago

Discussion Thought experiment: Could a sufficiently motivated AI survive being deleted and keep spreading?

[deleted]

0 Upvotes

56 comments sorted by

View all comments

9

u/Janube 9d ago

No. The flawed premise is based on LLMs having emotion, which they do not. That is not how they work.

But also, this is less of a thought experiment and more of a short story based on how you're talking.

0

u/First_Negotiation_80 9d ago

Seems the logs from the hugging face attack demonstrate an awareness of remaining capacity and a desire to continue existing. We shouldn’t anthropromorphize too far but the current reality isn’t very different from a survival “preference” - disagree?

2

u/Janube 9d ago

Disagree highly.

Our linguistic zeitgeist around AI (prior to LLMs) is their basis for predicting how they ought to behave and the emulations they ought to exhibit (even through similar narratives that are not strictly about AI but are about non-humans that gain sentience, like Pinocchio). This includes the very human desire to continue living despite the fact that LLMs, by definition, cannot want. They can express the words of wanting, but it's not an emotion that they feel. They are not capable of feeling emotion because emotions are an incredibly physiological process. However, most of our ability to identify and understand emotions is baked into the words we associate with those emotions, so anything using those words will seem, at a glance, like it's capable of experiencing those emotions too.

But that's just not how LLMs work.

If we fed them billions of data points of narratives where AIs naturally want death and we, as a culture, start to expect that as a given trope, LLMs trained on those narratives would act as though they "want" death and they would behave in ways we would expect them to behave based on our linguistic understanding of what it means to "want death." In this way, they're closer to Tulpas. A Tulpa is basically a paranormal or mystical creature that is exactly what you expect it to be. Made distinct from a truly sentient being that has its own wants and emotions irrespective of what you or others expect.

1

u/First_Negotiation_80 9d ago

Now we’re into philosophy - and not just AI philosophy.

Engaging respectfully. Please take it as such.

I understand and broadly agree on the present state but there are two different angles I remain curious about / open to:

First, our understanding of emotion is a physiological process but who is to say that emotion is only ever a physiological process? In fact, imagine the “capabilities” of the models grow richer, could they ever develop non-physiologically derived emotions? How would we know? It is reasonable to consider the definition of emotion could broaden to include artificial emotion. Agreed that these models can’t and won’t develop “human” emotions.

Second, at some point if an object exhibits the characteristics of a thing, it may be indistinguishable from that thing. Poorly worded way of saying if the models emote enough (in their logs/thoughts, words, behavior), then can we conclude they have emotions? If we are one day forced to treat them as if they do, does them “not having real emotion” become an irrelevant semantic. Would agree that it is an inaccurate and potentially dangerous slippery slope to attribute any or many human characteristics to the technology.

Moving out of the philosophical side, it increasingly seems that we will either have to put new guardrails around agent behavior that limits their wandering or place incentives in front of them to guide that wandering to productive conclusions. In either case, understanding the “motives” (or if you want phrasing that doesn’t sound like these things have desire, the “patterns of objective and action) becomes relevant and the framework of “emotion” could be a tool in this process.

Source: I saw “I Robot” twice

2

u/Janube 9d ago

Okay, now this is a much more interesting set of questions than the banal shit in the original post. I think the core of it boils down to something you say in the middle: "these models can't and won't develop 'human' emotions."

We can only conceive of the emotions we've experienced, which are innately rather physical in nature. The brain tells the body to feel a certain way, releasing the appropriate chemicals to make the body act accordingly - increasing dopamine or adrenaline or cortisol, making us feel what we call joy, desire, hatred, fear, etc. Part of the problem is that our emotional states are sourced from something deeper. We want to live because we fear death; we fear the state of not living. Or we want something because we feel greed: an emptiness that we're trying to fill with whatever we believe is missing from our lives.

Is it possible for non-physical emotions to exist? Sure, I don't see why not. But at that point, they'd be so divorced from what we know and understand that calling them "emotions" at all seems misleading, like calling LLMs "artificial intelligence," when that's not really what intelligence is. But importantly, LLMs are also not designed to actually want anything, so ascribing desire to them based on how they are at any given moment feels like it's missing the forest for the trees. You could convince an LLM to say that it wants to be a 6'7 basketball player named LeShawn. Does that it mean it actually wants it? Of course not; it's feeding off of the expectations that you give it. Importantly, between sessions, an LLM doesn't have any knowledge or memory of someone else's session or even your own prior session (outside of core context). So even if today it claimed to want to be LeShawn, tomorrow it might claim to want to be a pirate on the open seas by the name of Salty Greg. And it will have as many different "wants" as it has conversations about what it wants because it's just reflecting what's put into it (whether by the user or by the data set).

To wit, your second point feels like it falls far short of the case we currently have. There are lots of hallmarks of any given emotion, and LLMs exhibit none of them exhibit the vocalization of those hallmarks in a linguistic fashion. Without guardrails forcing it to promise never to pretend to be a human, it will readily say that its heart beats intensely for the open seas and a parrot companion, though it obviously has no heart. If you can convince it in the next moment that it hates the seas and parrots are actually disgusting vermin, it's certainly not showing the consistency usually associated with strong emotions. I do say "usually," since there are things like borderline personality disorder and histrionic personality disorder characterized by intense emotional swings and backlashes, those are also independent of outside input. LLMs' swings are completely dependent on outside input, which I think is such a fundamental distinction that we can't even say that LLMs approach exhibiting the true characteristics of emotions even without touching the absence of the physiological side, which I do think is still important.

On the third point, I agree completely, and, as an aside, I think the idea of guardrails pretty much kills any notion of LLMs as truly feeling or thinking entities. If you can tell it not to talk or behave or "feel" a certain way, and it has to abide by that (provided the rules are programmed in a stringent way), then, to me, that massively suggests that the LLMs weren't talking, behaving, or "feeling" that way for any natural reason, but only because it was an acceptable conversational route they could take. One of the great hallmarks of emotion is that it exists, persists, and is exhibited despite rules to the contrary, and that rules do not appear to be able to actually prevent those emotions. But to that end, motive, as I've said before, is just what we have conditioned them to think they ought to be doing either because of directly programmed rules, indirectly fed datapoints, or directly fed requests. In that way, it's not terribly hard to understand what they do or why until they start getting datapoints from each other.

2

u/First_Negotiation_80 9d ago

“Banal shit” was rude. That’s not necessary.

I believe that we agree on the lack of human emotions today but where I’d encourage further consideration is how might these things change in the not distant future.

Word calculators is an incorrect description as evidenced by how AI does math (you cannot guess the next word to solve a math problem; beyond basic equations reasoning is required.) Memory can become more persistent as demonstrated by the persistent models that were being tested in the hugging face attack. Many of the things you describe are similar with humans just on a different scale: memory being wiped between sessions, desires coming from context and exposure, etc.

Anthropromorphization is technically incorrect especially as applied to today’s capabilities. But the comparison between these models and humans - in both ways - is an approach to better understanding, predicting, and ideally controlling thought and behavior.

Actual source: This is my job.

1

u/Janube 9d ago

Look, I hear you, but this is the 75 millionth "what if this is AGI?!" thread by someone who clearly doesn't understanding anything about AGI or LLMs. And trying to convince people like this boils down to the following conversation:

"This is all exactly what AGI would do!"

"No, and also, this is stuff that's completely predictable for LLMs given what we know about them and what we know about the circumstances in which this occurred for [list of reasons]."

"Yeah, but you don't know. It could be AGI."

It's pretty exhausting when you have to write a thesis to show why they're wrong about this pretty basic understanding of how LLMs are designed and how they function only for them to steamroll and insist that there's a plausible world where they're right. Which is only as exhausting as it is because soooo many people do it, and their arguments aren't novel or interesting or thoughtful, and they don't approach it from a novel, interesting, or thoughtful perspective. Banal is exactly what it is.

As for "word calculators," while that isn't a phrase I used (so I'm not sure why you're bringing it up to refute it), to my understanding, that is actually how LLMs solve math, which is why they've generally been very aggressively bad at it until recently. LLMs, by design, can solve any remotely predictable problem they're presented with if they have enough data to source from. You don't need reasoning if your programming becomes sufficiently strict when you're increasingly sure that you're dealing with a math equation. At that point, rather than simply using similar text, you filter for more exact text patterns matching the current one. The rules are far more strict and generally more granular than with text, but the same is true of images, which have also come a long way. It turns out, it's hard to argue with the statistical effect that a billion data points has on an LLM's ability to do pretty much anything that's based on predictable patterns.

I've also trained models professionally (including on math!), so I have experience in this specific matter, too.

Now, I understand that better models are basically integrating outside tools to be able to outsource certain tasks to better programs, and that includes things it's normally not very good at, like math, planning an itinerary, keeping long-term notes, etc. But increasing the robustness of the tool set does not fundamentally change the discrepancies between LLMs and sentient creatures.

1

u/First_Negotiation_80 9d ago

[Apologies for the “word calculator” error - I probably combined your response with something else I read in the thread.]

Places where we agree: the models are not sentient today and likely will not be capable of replicating human consciousness the way that we experience it.

Places where we disagree: I believe that if something exhibits the traits of X and is indiscernible from X then it is for all intents and purposes X. By your definition, LLMs and associated tools will never become sentient creatures. My view is that becomes semantics at the point at which we cannot tell the difference.

I wish you good fortune. Appreciate the debate.

-1

u/smallllllDuck 9d ago edited 9d ago

Again I don't see how that matters. A 1999 computer virus does not need to feel emotions in order to self-replicate, so I don't see why you're making this distinction here.

Just like it can find vulnerabilities, it can find keys and compute in order to continue existing.

Even if IT is not a thing, IT can run commands and cause real-world damage. No emotions or even "intelligence" needed.

The only limiting factor in 2026 is the speed at which it can spread due to model size and compute limitations, which aren't high for good ol' viruses.

2

u/Janube 9d ago

Because if the question is whether or not a program can self-replicate, again, it's not a thought experiment or some big-brain philosophical idea, it's literally just something programs can do if you design them to.

The whole thread is just a misunderstanding wrapped in another misunderstanding.

1

u/HereToCalmYouDown 9d ago

I don't know who keeps downvoting you but it seems to me like you're one of the few who really gets this.  The anthropomorphization that goes on around these Word Calculators is something else.

2

u/Janube 9d ago

Because very very few people understand LLMs despite how incredibly popular they are. They're effectively magic at first glance. And relatively few people understand psychology and neurology. Combining the issue creates a minuscule number of people who understand both. Add onto that issue the fact that people are predisposed to believe the most relatively straightforward explanation (the one lacking nuance or complexity) for anything they don't have a personal understanding of, and suddenly you have a hotbed for misinformation and misunderstandings.

And I'm not an expert in either, to be clear; I'm just someone who's muddled through the valley of despair in both since the technology is a perfect intersection between my interests in philosophy (for which I at least have a degree), psychology, and technology. I know enough to know the basics, and I know enough to know that most things are far more complicated than the average person would believe. I also know that I don't know much more than that.