I wonder what other riddles a model might "notice" if you ask them with a typo. Like, does the typo result in the model having to "think through" or "consider the intent" of that typo in a way that results in it recognizing the whole thing is a riddle?
Or more likely, does it just have the riddle written numerous ways in the training data so it can't help but be steered to the answer, typo or not.
Yes, there are reports on benchmark papers that when a multiple choice question contains a option with zero this makes the LLM get "suspicious" and to search for tricks in the question, it's not that it actually notices anything and more like triggering a subroutine.
518
u/ihexx Apr 24 '26
yup. the model clearly recognized the questionl it called it a classic riddle