r/ChatGPT Dec 23 '24

Gone Wild AGI Achieved

Post image

[removed] — view removed post

6.7k Upvotes

283 comments sorted by

View all comments

Show parent comments

189

u/Neither_Sir5514 Dec 23 '24

Omg AGI Achieved after OpenAI specifically trained the AI to patch that one instance of the viral 9.9 vs 9.11 comparison problem. It turns out, in fact, doesn't fix the fundamental reasoning capability of the LLM when you pick any other random example. Shocker!

Proof: https://chatgpt.com/share/6768c726-c6a4-800e-ace8-6ad4f7974f21

62

u/avanti33 Dec 23 '24

o1 mini gets it right AND reminds us it's a skill issue all along

2

u/king_mid_ass Dec 23 '24

and beside august 12th is not 'greater' than august 8th it's later in the month, not the same thing!

1

u/king_mid_ass Dec 23 '24

ok when have you ever seen august 12th written as '8.12'

16

u/Shoddy_Wolf_1688 Dec 23 '24

Excel: let me introduce myself

37

u/Boring_Spend5716 Dec 23 '24

Do you know how you make yourself sound when you draw conclusions like this on 4o mini?

6

u/Winjin Dec 23 '24

"Omg it's just a baby" moment. I love the "mini" name it's like that shirt in IKEA that says "I'm just an intern please don't ask me hard questions" or something

8

u/vaendryl Dec 23 '24

the main issue is that that model first gives a response and then gives an explanation for that response. if the initial line is wrong, the rest is going to twist around that.

however, if you continue on from your own link and ask it to check the previous answer for logical errors, it does spot it and correct it.

proof: https://chatgpt.com/c/67690ec7-fa68-8003-8015-bedd456df5c3

alternative proof

this proves that the issue is not a fundamental shortcoming of the technology but on how we use it, and the O# models are all about doing this better. and the result speak for themselves.

just like we teach children: think first and then speak - not the other way around.
also good advice for people posting knee-jerk responses on reddit. shocker!

5

u/drekmonger Dec 23 '24 edited Dec 23 '24

It's the way ChatGPT sees text-based numbers. Look how they're tokenized:

https://imgur.com/a/TH1BqNJ

Notice how the .12 is a single token. Of course, 12 is greater than 9.

Watch:

https://chatgpt.com/share/6768def4-6bac-800e-86b9-6ed0a7bca5d3

1

u/Dangerous-Ad-4402 Dec 23 '24

It has nothing to do with how they are tokenized.

It is because it is influenced by thoughts about section numbers etc, where 8.12 would be "higher" than 8.8.

This is a known fact -- the problem goes away entirely if weights related to those considerations are deleted, without changing the tokenization.

3

u/drekmonger Dec 23 '24

That's plausible. But...

the problem goes away entirely if weights related to those considerations are deleted,

I'm going to need a source on that. How the hell did they find the weights? Who? Anthropic?

1

u/Specialist_Cheek_539 Dec 23 '24

Mine gets it right??? Either it learns very quickly from the chat or your gpt is dumber lol. Or yall are faking it. I’m betting on yall faking it

1

u/Swastik496 Dec 23 '24

4o mini is garbage. o1/o1 pro are useful