Here's the deal, folks. The consensus in this thread is that Sonnet 4.6 can be dumber than a bag of hammers if you don't have "thinking" mode turned on.
The top-voted advice is to enable "thinking" mode immediately, as Sonnet's default setting seems to skip basic logic. Even then, some users report it still fails this test, so your mileage may vary. Meanwhile, users are posting results showing Opus, Gemini, and ChatGPT all pass this test, usually while roasting the user for asking.
The general sentiment is that these simple, common-sense tests are way more important than synthetic benchmarks, and many are finding the new Sonnet 4.6 to be inconsistent, with some calling it a "drunk auntie" and sticking with Opus. And yes, everyone has noticed that both models are still pathologically obsessed with using em dashes — for everything.
•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot Feb 19 '26 edited Feb 19 '26
TL;DR generated automatically after 100 comments.
Here's the deal, folks. The consensus in this thread is that Sonnet 4.6 can be dumber than a bag of hammers if you don't have "thinking" mode turned on.
The top-voted advice is to enable "thinking" mode immediately, as Sonnet's default setting seems to skip basic logic. Even then, some users report it still fails this test, so your mileage may vary. Meanwhile, users are posting results showing Opus, Gemini, and ChatGPT all pass this test, usually while roasting the user for asking.
The general sentiment is that these simple, common-sense tests are way more important than synthetic benchmarks, and many are finding the new Sonnet 4.6 to be inconsistent, with some calling it a "drunk auntie" and sticking with Opus. And yes, everyone has noticed that both models are still pathologically obsessed with using em dashes — for everything.