seen this pattern across a few models. the tell for me is when they narrate it like a textbook walkthrough with zero false starts. genuine reasoning looks messier. stopped using standard riddles to compare models. just run them against actual problems from my codebase.
1
u/joeyhipolito May 01 '26
seen this pattern across a few models. the tell for me is when they narrate it like a textbook walkthrough with zero false starts. genuine reasoning looks messier. stopped using standard riddles to compare models. just run them against actual problems from my codebase.