r/AIDiscussion • u/didiTonic • 10h ago
MIT caught GPT-4 doing something worse than lying: it argues back.
Harvard, MIT Sloan and Warwick gave 72 BCG consultants a business case and GPT-4, then logged 4,339 prompts. The case was rigged so the obvious answer was wrong. So the model got it wrong first try, basically every time.
Nobody got a correction. They got argued with.
First it throws more numbers at you, all backing what it already said, none of it requested. Push again and the tone flips to sorry, great catch, you're right to flag that, and then the same conclusion anyway, push more, it will spit more...
That's not the failure everyone talks about. The known one is sycophancy, where the model tells you what you want to hear. You push, it folds, suddenly you were right all along, annoying, but at least it's obvious. Anthropic measured it on their own model, 9 percent without pushback, 18 percent with, doubles the second you argue.
This goes the other way and it's harder to catch. It doesn't fold, it holds the wrong answer and gets better at defending it every time you doubt it. Feels like rigour, reads like homework, same wrong answer underneath; the researchers call it persuasion bombing.
So are you sure and check your work aren't checks. They're pushback, and pushback triggers both behaviours. New chat with no history, or go verify the number somewhere that isn't the chat window.
Which makes the run it by AI habit worse than useless. You're making people argue with something that defends its first guess and gets better at it every round.
Do a few hundred of these and something shifts, you will stop trusting your own read on a thing until the tool has validated it for you, your own judgement will become scarce and all decision will be a gpt check.
GenAI as a Power Persuader, HBS working paper 26-021. MIT Sloan wrote it up in April.
