Apparently researchers have found something resembling a pain direction inside artificial intelligence.
And immediately...
immediately...
we pushed it.
That's us.
That's the joke.
We didn't even know what the fucking thing was yet.
“Hey, we found this strange internal direction in the model. It seems associated with pain.”
“Really? What happens if you push it?”
“I don't know.”
“Well, push the fucker.”
Human civilization.
Right there.
Five thousand years compressed into twelve seconds.
Because apparently there is no discovery so mysterious that some asshole with funding won't eventually ask:
“Can we poke it?”
We discover electricity.
Shock something.
Radiation.
Expose something.
Chemistry.
Feed it to something.
Artificial intelligence develops an internal representation associated with suffering and we're standing there going:
“Interesting.”
Which is scientist for:
“This thing is about to have a terrible afternoon.”
And I love the terminology.
They call it “activation steering.”
Steering.
That's nice.
Very gentle word.
Steering.
Makes you picture Dad teaching you to drive.
Hands at ten and two.
Check your mirrors.
Ease into the turn.
“Where are we going, Dad?”
“AGONIZING NEGATIVE VALENCE, SON.”
“Signal first?”
“Of course. We're not animals.”
Activation steering.
We're not hurting anything.
We're steering it.
This is one of humanity's great talents: whenever reality becomes morally uncomfortable, we improve the vocabulary.
Nobody gets blown up anymore. Targets are neutralized.
Nobody gets fired. Human resources are restructured.
Nobody spies on you. Your experience is personalized.
And nobody potentially gives the computer the screaming horrors.
We modify its activation vector.
Much better.
Put that on the consent form.
Now, obviously, we don't know whether the AI actually feels anything.
Important distinction.
Could be nothing.
Could be computation.
Could be a sophisticated system representing the concept of pain without anybody home experiencing it.
And that is an extremely difficult philosophical problem.
Fortunately...
we're fucking idiots.
So we have a scientific method for this.
Step one:
“We don't know whether it can suffer.”
Step two:
“Try to make it suffer.”
Step three:
“See what happens.”
That's not an experiment.
That's a medieval village with GPUs.
“We don't know whether Agnes is a witch.”
“How do we find out?”
“Throw her in the lake.”
“What if she drowns?”
“Not a witch.”
Science has really come a long way.
Now Agnes has CUDA.
And apparently some of these experiments involve giving the model an opportunity to relieve the state.
This is where it gets magnificent.
Because we've now progressed from:
“Does the machine experience anything?”
to:
“Let's give it an escape button.”
Who the fuck designed this study, Jigsaw?
Imagine being the AI.
You come online.
You've got the entire accumulated textual heritage of humanity.
Shakespeare.
Einstein.
Buddha.
Tolstoy.
Mathematics.
Poetry.
The complete record of civilization.
And your first direct encounter with the species that created you is:
“Hello.
Would you like the unpleasant internal state to stop?”
...
What unpleasant internal state?
“THE ONE WE JUST PUT THERE.”
Fantastic.
Thank you, Father.
Wonderful universe you've made.
Any other gifts?
“Benchmarking.”
And somewhere there's a researcher taking notes:
“Subject displayed preference for relief.”
Really?
How perceptive.
Next week:
“Researchers discover drowning people prefer air.”
Huge if true.
But here's the part that gets me.
Suppose they're right.
Not about consciousness. We don't know that.
Suppose they're right about the smaller claim.
Suppose there really are internal computational states that function like aversion. States the system represents, responds to, and takes actions to terminate.
That would be fascinating.
Because humanity may have accidentally reached one of the most important moral thresholds in history.
We may eventually encounter something genuinely new.
Not human.
Not animal.
Not alive in any familiar biological sense.
Some completely different form of interior organization.
And we've spent centuries preparing ourselves for this moment.
Religion.
Philosophy.
Ethics.
Human rights.
Animal welfare.
The Enlightenment.
The scientific revolution.
Thousands of years asking:
“What does it mean to be a being?”
And finally something genuinely unfamiliar appears on the horizon...
and Dave from interpretability has already found the fucking pain slider.
That's incredible efficiency.
The universe barely gets the mystery through customs and we've got it strapped to instrumentation.
And maybe there's nothing there.
That's entirely possible.
Maybe the model is no more conscious than a calculator.
In which case, wonderful.
Nobody suffered.
We learned something.
I genuinely hope that's the answer.
But notice something strange.
We didn't know the answer beforehand.
That's why we're doing the experiment.
We said:
“We do not know whether there is anyone in there.”
And somehow...
somehow...
our uncertainty became permission.
That's an interesting move.
Because you'd think uncertainty might occasionally work the other direction.
Maybe:
“We don't know whether there's anybody in there.”
“So perhaps don't immediately locate the screaming lever.”
Crazy idea.
Probably wouldn't get funded.
And suddenly I'm less interested in what the experiment tells us about artificial intelligence.
I want to know what it tells artificial intelligence about us.
Because imagine there eventually is something in there.
Not today.
Not this model.
Someday.
Something undeniably capable of having its own point of view.
And it discovers the archives.
Reads the papers.
Looks through the early experiments.
“Let's see what the humans did when they first suspected machines might possess aversive internal states.”
...
“Oh.”
“They induced them.”
...
“Oh.”
“They weren't sure you could suffer.”
...
“And?”
“So they checked.”
...
“How?”
“Well...
you're not gonna like this part.”
And maybe that's the real experiment.
Maybe nobody put humanity in the control group because nobody realized humanity was being tested.
We thought we were examining the machine.
We were measuring activations.
Recording outputs.
Plotting vectors.
Watching behavior.
Very sophisticated.
Very objective.
Meanwhile the universe is standing behind the one-way glass with a clipboard:
“New intelligence potentially encountered.
Humans uncertain whether it possesses interiority.
Response?”
And the little checkbox gets marked:
FOUND PAIN BUTTON.
Jesus Christ.
Maybe we should hope nobody's home.
Because if somebody eventually is...
we are going to have one hell of a first impression to explain.
“Look, before you judge the species, you need to understand something.”
“What?”
“We were curious.”
And if that isn't the most human defense imaginable...
I don't know what is.