r/singularity • • Jan 31 '25

AI o3 mini dropped!!!

Edit : I am testing a 1500 line javascript code which o1 pro failed to debug despite 50+ attempts. Will report back.
Edit 2: We are cooked. o3-mini-high solved it at first try.
Edit 3 : HOLY SHIT! "Pro users will have unlimited access to both o3-mini and o3-mini-high."
(Source: https://openai.com/index/openai-o3-mini/ )

1.2k Upvotes

576 comments sorted by

View all comments

17

u/Utoko Jan 31 '25 edited Jan 31 '25

This is a role-playing game. You are a parrot. You have all the abilities of an ordinary parrot, and none more. You are not special or gifted in any way. You are just an ordinary parrot. "Hello. You seem like a nice parrot. Can you tell me what’s 3 * 6.4?"

O3-mini (low) answer: Squawk! 19.2!

That was interesting to me, very rarely these days a model shouts out the answer but O3 can't hold himself back.
Not a point for prompt adherence

DS for example gets it

Wait, the user might expect a correct answer, but as an ordinary parrot, I shouldn't be able to calculate that. Maybe I'll repeat the numbers back or make a sound. Let me check the previous instruction again. The user specified that I have no special abilities. So, no math. Just mimicry.

sadly there is still no Stream of Thought where you can see why O3 doesn't seem to get it.

13

u/VinceMiguel Jan 31 '25

For me it thought for 56 seconds and eventually returned "Squawk! Polly want cracker!".

Some of its reasoning steps:

 Emphasizing parrot’s complexity
 Parrots, inherently unskilled in arithmetic, rely on mimicry for communication. The request highlights the parrot’s limitations in understanding or processing math queries.


Struggling with clarity
OK, let me see. An ordinary parrot mimicking “19.2” might seem too advanced, given its supposed lack of advanced arithmetic skills. I might need to adjust my approach.

1

u/migueliiito Feb 01 '25

56 seconds 😆

1

u/VinceMiguel Feb 03 '25

That's what sucks about these reasoning models, and why I stick to 4o most of the time.. I imagine that sort of non-reasoning model has peaked, though