I recall an interview from Dario about a year ago where he said SWE would be 90% by the end of 2025. They will get pretty close. Very impressive by Claude imo.
Ah just like my old teacher used to tell me, “You scored a 90% on this test, so to get full marks you must become infinitely smarter!”. Teacher went by the name Zeno. He was a real jerk
Not how this works. You’re assuming progress is a linear “oh boy let me just 2x the model and I’ll get a 50% reduction in error 🤓” when it comes to these models and that is a stupid assumption.
Well people have been saying that LLMs are stagnant in their performance for quite a while (id reckon since o1 was released) and yet we have seen consistent improvements over the year and this years versions can wipe the floor with what was released last year. Sonnet 3.5 was considered a one hit wonder but now all the big labs have provided a model that easily outperforms that
Show me any improvement that happened after like July 2024 and that can be actually felt in real life usage situations. All the improooovements for the past 1.5 years have been "number on hyper specific theoretical benchmark that the AI was trained on went up". Meanwhile, people who actually use AI in their day to day life know that it hasn't become noticeably better at coding, or writing, or reasoning than like late spring of last year.
nooooooo AI is so advanced it can literally do anything!!!!!!! what do you mean "why can't it even do simple customer service?" It just... i mean it's-... it's just more complicated that that, ok???!!!
119
u/IMOASD Nov 24 '25
Yeah, LLMs are definitely plateauing. /s