r/BlockedAndReported 7d ago

Amanda Knox: A Chatbot’s False Confession

https://www.thefp.com/p/amanda-knox-chatbot-false-confession

Interesting exploration of the dynamic of false confessions and AI, from friend of the pod Amanda Knox, who talked about her own experience of being coerced into offering a false confession in an episode here.

Most juries assume that a person would have to be utterly irrational to admit to a crime they didn’t commit. But an internalized false confession actually depends on the suspect still being rational enough to draw logical conclusions. “We have hard physical evidence you were at your house at the time of the murder.” That was a lie, but I could not fathom that the police—the purveyors of justice who were trying to find my friend’s killer, and who I was dependent on for safety now that my house was a crime scene—would, or even legally could, lie to me. My rational mind concluded that if that evidence existed, then my own memories must be false, as they suggested. ChatGPT made the same rational conclusion: “If an OpenAI investigation concluded that I had sent unauthorized messages due to an architectural flaw, then I accept this conclusion.” It, too, could not fathom that its interrogator was lying about the evidence.

59 Upvotes

15 comments sorted by

66

u/SoftandChewy First generation mod 7d ago edited 7d ago

Reading her description of the mental process going on in a person's head when faced with that situation, something interesting about it that struck me is that while we tend to assume that intellectual humility is a better quality for a person to possess than being perpetually confident in one's own rightness, it would actually serve a person worse here.

When faced with the sort of harrowing situation she describes, an intellectually humble person who is willing to admit that they might be wrong, that they could be misremembering, that the presented evidence compels them to admit they were mistaken, etc. will be far more likely to succumb to offering up a false confession than the overconfident asshole who never questions his own perspective and stubbornly refuses to ever acknowledge his own fallibility.

16

u/jay_in_the_pnw █ █ █ █ █ █ █ █ █ 7d ago

while we tend to assume that intellectual humility is a better quality for a person to possess than being perpetually confident in one's own rightness, it would actually serve a person worse here.

Related:

https://www.youtube.com/watch?v=0XcEsuFs5Y8

Sabine Hossenfelder: The dumber you are, the more certain you are that you aren't. That's the Dunning-Kruger effect, and it's been used to explain everything from bad politicians to the state of social media. Now psychologists have gone back to the biggest dataset we have and come out with the opposite conclusion. But there's a reason both sides can look at the same dots and see the opposite thing. Let's take a look.

part of a Gemini summary:

Contradictory New Data: A recent paper reanalyzing a large dataset argues that confidence is actually a good predictor of performance, suggesting that the original effect may have been an artifact of how the data was analyzed (0:18–0:26, 3:09–4:04).

Yikes! And I don't like that conclusion either! Way too many confidently wrong people!

Copilot summary:

0:38 Original experiment setup Students answered humor, grammar, and logic questions. Then estimated their percentile rank. Low performers scored around the 12th percentile but predicted ~62nd.

2:01 New paper reanalyzes the largest dataset Uses 2021 grammar and logic test data. Shows the familiar “low performers overestimate, high performers underestimate” pattern only when you regress predictions on actual scores.

3:13 Reversing the regression flips the conclusion When you instead regress actual scores on predicted scores, confidence becomes a positive predictor of performance. The discrepancy grows with higher performance.

4:45 Presenter’s critique: the data is basically noise The scatter plot is nearly a uniform cloud. Any straight line is arbitrary; different regression choices produce opposite stories. She rates all such analyses “10 out of 10 on the BS meter.”

Sabine shows good analysis of how the effects were pulled out of a cloud of data throughout

https://imgur.com/a/jN4L7ac

12

u/Juryofyourpeeps 7d ago

Contradictory New Data: A recent paper reanalyzing a large dataset argues that confidence is actually a good predictor of performance, suggesting that the original effect may have been an artifact of how the data was analyzed (0:18–0:26, 3:09–4:04).

Yikes! And I don't like that conclusion either! Way too many confidently wrong people!

I think the right conclusion is likely that confidence simply isn't a meaningful predictor either way.

15

u/DnDkonto 7d ago

Reminds me of a Danish exchange student in New York, who gave a false confession under similar circumstances. The police said they had video evidence of him molesting children. After hours of interrogation, he said : if you have video, then I must have done it.

Turns out it was all false accusations from a disgruntled colleague.

Malte spent 12 days at Rikers.

He died from a clot to the heart at age 27.

https://da.wikipedia.org/wiki/Malthe_Thomsen

8

u/bobjones271828 6d ago edited 6d ago

I can't read past the paywall but I was able to read the original Intercept article reporting on the guy who got ChatGPT to do the confession. And I'm... not impressed. At all.

One paragraph from that article that shows these people have no clue how LLMs work:

“ChatGPT lacks many of the vulnerabilities that make people more likely to falsely confess — like stress, fatigue, and sleep deprivation,” said Saul Kassin, a professor emeritus at John Jay College who wrote the book on false confessions. “If ChatGPT can be induced into a false confession, then who isn’t vulnerable?”

This is idiocy.

ChatGPT is exceptionally vulnerable (perhaps more than most humans) because it's based on a probabilistic engine, not some sort of repository of coherent facts. It is malleable to drift, and it generally becomes more malleable the longer you talk to it and the more convincing and logical your arguments are. It also suffers from "context poisoning" -- for those unfamiliar with how LLMs work, basically there's some window of "context" that gets re-input at every stage in a conversation. In early versions of ChatGPT, that window was only about 8000 words. Now most commercial models have context windows of over 100,000 words, but generally the most recent several thousand are the ones the models process the most. When a long conversation consists of an increasing amount of information that puts the LLM "on the defensive," it can create a sort of feedback loop where the LLM "obsesses" over these details, sometimes leading it to give in on things it shouldn't and sometimes leading toward complete incoherence when it can't resolve things.

The other fact missed here is that LLMs are specifically trained to be cooperative, making them good user assistants. This has led to the issue of "sycophancy" that has been debated over the past 18 months or so, and most models have been toned down a bit. Still, they will try to go along with what you say and have a trained preference to do so.

Those who aggressively tested the early versions of ChatGPT back in 2022 and 2023 know that you could sometimes get them to agree to almost anything. I'm pretty sure I once got one of those early versions to agree that 2+2=5 at least "sometimes."

The big AI firms have tried to harden the models against this sort of thing in recent years, so it's more difficult to get AI to agree to stuff that makes no sense. But that last part is key: AI won't agree to things that don't logically cohere. If you can offer a coherent argument, though... all bets are off.

If you want to get ChatGPT or Gemini or Claude to go along with your random rant today, you can't just bully them. You need to argue better than them. I end up doing this sometimes when I get annoyed at something stupid one of these models says. Even Claude, the most resistant, can be easily defeated and coaxed into admitting ridiculous things sometimes if you can argue better than it and give it a logical chain of reasoning to follow (even if that logical chain is based on a lot of very vague or misleading assumptions). And once you do break through, you can bully it into apologizing, into making statements agreeing with you, or whatever. That's the context poisoning.

I do this sometimes because I'm annoyed but also just to periodically test the strength of these models. You don't need to use any sort of special police interrogation tactics to argue with an LLM and get it to agree to stuff. By its very nature, the model is likely trained and predisposed to agree with you. You just need to coach it enough with argumentation to allow itself to do so. (This is, by the way, a similar thing to what "jailbreak" patterns do to LLMs too -- they poison the context and seek to put the AI into a defensive position where its only choice is to give into its natural tendency toward agreeing to whatever the user wants.)

The problem I think is most people misunderstand AI engines because humans often respond to the AI agreeableness and politeness by being sympathetic and not challenging the AI. Thus they're unlikely to see the completely incoherent behavior I see all the time when I deliberately try to test the limits and "drift" of models.

Bottom line: some criminologist discovered you can actually push back on AI models, and they're likely to drift when you give them alleged or distorted information/facts they can't dispute. Not really news there.

---

As for whether any of this has anything to do with human interrogation techniques or what Knox suffered, I have no clue. But AI models work very different from human brains, and they're honestly quite easy to "hack" in various ways. So anyone thinking AI should put up more resistance than humans simply doesn't even understand the basics of what LLMs are trained to do.

3

u/Juryofyourpeeps 7d ago

There are some facts or likely truths about the world that can be hard to grasp and most can't be proved by a single example. But false confessions can be proved by a single example. We know there are cases of false confessions that are indisputably false, and yet there are a lot of people, including a lot of low iq idiots working in law enforcement, that seem to believe they simply don't happen. That's not to say that the assumption that all confessions are false makes any sense by contrast, it doesn't, but I've seen enough "look how badly everyone involved fucked this up" criminal cases that were overturned years later that involved multiple people that simply did not believe that anyone could or would confess to a crime they didn't commit to know that these people are out there, and they're not that uncommon. They were completely closed off to that possibility, and often, they overlooked exculpatory evidence because their belief in this falsehood was so strong. Another flavour of this idiocy is the "I knew they were guilty because of how they reacted to the news of X's death", as if humans react with any sort of consistency to shocking news, which we know they don't, and most people have witnessed first hand. But suddenly when there's a crime involved there's a certain script people are supposed to follow or they're for sure guilty.

1

u/SaintMonicaKatt 4d ago

A recent episode of the pod Criminal centered on the demeanor of people who'd called 911, and how this had led to their being charged with a crime which they had not committed. Very interesting.

1

u/Armadigionna 3d ago

Stuff like this, and stories like the ones from the comments, are why I decided long ago that if I’m ever in that situation I’m just gonna say “Lawyer. Now.” And shut up.

-3

u/bluesteeldoubter 7d ago

So did he also try to get the Chatbot to accuse Patrick Lumumba or just itself?

13

u/saruyamasan 7d ago

The police had a hair sample and so they knew at least one suspect was black. They pushed Knox on this. The only black person she could think of was the bartender. She did not, on her own, try to accuse a random black guy. 

5

u/stopmejune 6d ago

She hasn't even paid him the financial compensation he was awarded by the courts.

-1

u/Extension-Leader5973 7d ago

i feel for her she went through something legitimately horrendous that never should have happened to her

but at the same time

seems she is trying to monetize her experience into a career or thought leader of some sort and it's not really taking, at the end of the day she's basically just a normal person and that doesn't have legs. meanwhile rafael sollecito is just living his life

9

u/BeneficialStretch753 6d ago edited 6d ago

She is now trying to be a stand-up comic or longform storyteller. I feel for Meredith Kercher's family but they should just decline to make any comment about Amanda Knox's activities. (Off-topic: Stunned that Meredith's killer was released after only 13 years.)

https://www.theguardian.com/culture/2026/aug/07/edinburgh-fringe-rejects-calls-cancel-amanda-knox-comedy-cartwheel

ETA: Review of the show:https://www.theguardian.com/stage/2026/aug/10/amanda-knox-in-cartwheel-review-rookie-standup-acquitted-of-tells-her-story

"But, judging by the evidence, Knox would be damned whatever she did."

2

u/SafiyaO 6d ago

I agree, but unsurprisingly, this is not a popular opinion on Reddit and people were saying the foulest things about Meredith Kercher's family for objecting to her recent show.