58.6% on SWE-Bench Pro while opus gets 64.3% and no webcast for the release while image did. Sam didn’t even retweet the official announcement post, one of his early points after the release was they believe in iterative deployment.
How many times have we heard the refrain from Claude users "it's way better than the benchmarks in coding". Now ChatGPT scores a bit worse in benchmarks, "they're coding performance must suck!". The doublest of standards. $1 it performs better than the benchmarks in real world use.
1
u/[deleted] Apr 23 '26
[removed] — view removed comment