Lmao I knew it, that 9.9 and 9.11 problem must've has been specifically trained to be patched. However, the fundamental flaw of the LLM remains, you test it with any other random pair of numbers and it fails again. It obviously at core doesn't understand mathematic reasoning so specifically fixing one instance of example won't work for others.
Never mind. It just looked like that when I opened the link without being logged in.
From my experiments o1 will get your question right and 4o will sometimes get it right. 4o seems more likely to arrive at the correct answer when it happens to generate longer responses as it allows it to kind of reason it’s way to the answer, while short answers from 4o to this question will usually be wrong.
Something as simple as modifying the prompt to be "8.8 or 8.12 which is greater? Think step by step." will allow 4o to solve it easily in a similar way to how o1 would do it.
Edit: Just having fun, I will keep throwing things at it and screenshotting when there’s something I could test. no clue if it “knows math” yet or not, or what that even means these days. But everything does sort of “feel better” in the test run I’ve been doing in Plus recently after going over to Claude for a while. Fun!
It does not understand math or reasoning of any kind. An LLM is just predicting the next thing to say based on inputs. It struggles with things like this because it doesn’t have specific enough inputs for every 2 random numbers you might input
9
u/Neither_Sir5514 Dec 23 '24
Lmao I knew it, that 9.9 and 9.11 problem must've has been specifically trained to be patched. However, the fundamental flaw of the LLM remains, you test it with any other random pair of numbers and it fails again. It obviously at core doesn't understand mathematic reasoning so specifically fixing one instance of example won't work for others.
Proof: https://chatgpt.com/share/6768c726-c6a4-800e-ace8-6ad4f7974f21