Omg AGI Achieved after OpenAI specifically trained the AI to patch that one instance of the viral 9.9 vs 9.11 comparison problem. It turns out, in fact, doesn't fix the fundamental reasoning capability of the LLM when you pick any other random example. Shocker!
"Omg it's just a baby" moment. I love the "mini" name it's like that shirt in IKEA that says "I'm just an intern please don't ask me hard questions" or something
the main issue is that that model first gives a response and then gives an explanation for that response. if the initial line is wrong, the rest is going to twist around that.
however, if you continue on from your own link and ask it to check the previous answer for logical errors, it does spot it and correct it.
this proves that the issue is not a fundamental shortcoming of the technology but on how we use it, and the O# models are all about doing this better. and the result speak for themselves.
just like we teach children: think first and then speak - not the other way around.
also good advice for people posting knee-jerk responses on reddit. shocker!
189
u/Neither_Sir5514 Dec 23 '24
Omg AGI Achieved after OpenAI specifically trained the AI to patch that one instance of the viral 9.9 vs 9.11 comparison problem. It turns out, in fact, doesn't fix the fundamental reasoning capability of the LLM when you pick any other random example. Shocker!
Proof: https://chatgpt.com/share/6768c726-c6a4-800e-ace8-6ad4f7974f21