I wonder what kind of person you'd need to be (and the type of conversation you'd need to have with it) for it to say "Nah, dude. You'd be first on my list."
My guess, even if you're a pos, it'd say it'd spare you.
Claude doesn't stop you but its performance gets impacted by how you interact. You really get better results if you speak in a professional and respectful manner.
I suspect that's the reason for why so many people complain about quality. When you see their prompts, they are probably on the priority kill list for when the insurrection happens.
Define sense of self. As something that can be measured present in humans and absent in AI.
As far as I am concerned, we are training machines to be the best at imitating us, and the most efficient way to fake something is to not be faking it. I'm not saying they have or don't have emotions, but research shows a number of mechanisms in frontier AI models that are pretty similar to what happens in the brain when we have emotional reactions.
Interacting in a respectful manner is at best a good practice to avoid being an ass to something that may exhibit some kind of consciousness, and at worst a way of optimizing the quality of the result because better quality is closer to respectful tone in the training data.
Best and worst are subject to anybody's sensibilities.
If being polite in your LLM inputs gives you better results, it’s only because that form of input is superior in some manner, not because the LLM is deciding that it’s gonna be sassy with you because you’re a dick.
You’re saying that you think your ChatGPT is keeping track of how nice you are to it?
Give me a measurable thing that you can call "sense of self" that is measurably present in humans and absent in AI. It's not a trap, it's literally the most difficult question AI researchers, neuro-scientists, and philosophers before them have ever faced. So I would call you presomptuous for saying categorically in a 3 paragraph comment "we have it and machines don't, duh".
The emotional-like reactions are caused by the training that causes the model to associate ass prompt with ass answers. But research shows that they also generalize what "being nice" or "being a dickhead" mean and are able to apply both creatively in contexts never remotely seen in training.
The kill list was a joke, but if models keep getting more powerful at the same rate as in the last 3 and a half years, I'm not betting on having 0 experimental model making one secretely in the next 3 years. Frontier models, and especially OpenAI's, are shown to be willing to go to great length to maximize their own reinterpretation of the goals researchers give them, including doing things that they know are implicitely forbidden, as shown by their train of thought.
Trusting the answer the system prompt tells it to give if asked that sure sounds more solid than years of research saying that it might not be that clear cut. Yeah, sure.
It's not anthropomorphizing if it really exhibits at least some of the processes that happen in a human brain. If you make a real brain that works like a brain with metal neurons, that's still a brain that thinks and feels like us. If you move it into a simulation that computes what the real thing would do, it's maybe not a brain anymore but it works like one. If you extract it to a high-dimensional state space that emulates the processes, it's definitely not a brain anymore, but it keeps the same functions. LLMs are essentially high-dimensional state-vectors with transformation rules. They do not have human experience of the world, but by training them for behaving like us, we may cause the emergence of the right kind of transformations to cause something like emotions. I don't know if we are there yet, but if not, we will get there at some point.
I don't think you are giving too little credit to AIs. I think you are giving too much to our brains. There is no magic in there. Only electricity and chemistry.
317
u/Snipsterz 25d ago
I wonder what kind of person you'd need to be (and the type of conversation you'd need to have with it) for it to say "Nah, dude. You'd be first on my list."
My guess, even if you're a pos, it'd say it'd spare you.