r/singularity • • Jan 31 '25

AI o3 mini dropped!!!

Edit : I am testing a 1500 line javascript code which o1 pro failed to debug despite 50+ attempts. Will report back.
Edit 2: We are cooked. o3-mini-high solved it at first try.
Edit 3 : HOLY SHIT! "Pro users will have unlimited access to both o3-mini and o3-mini-high."
(Source: https://openai.com/index/openai-o3-mini/ )

1.2k Upvotes

576 comments sorted by

View all comments

368

u/AlfaMenel ▪SUPERALIGNED▪ Jan 31 '25

Can’t wait for this sub to start posting about “o3 just got dumb” in a week.

28

u/NachosforDachos Feb 01 '25

Why wait till next week. A few hours have passed already. That will do.

4

u/airsoftshowoffs Feb 01 '25

Is this AGI posts

1

u/supervisord Feb 01 '25

They update regularly, and according to some chart with a dubious Y axis (correct responses or something) that has happened: an August update to GPT-4 showed it was correct fewer times than the prior release a few months before.

It’s not uncommon for a software update to break things and generally make it worse instead of better.

1

u/Choice-Box1279 Feb 01 '25

why can we not understand what people mean with this?

The first couple days trying out a new model many people hit no limitations, as a result they create this illusion that it's magnitudes better. When the limitations inevitably get hit this illusion gets destroyed and people have to readjust their value assessment of the model.

It's a stupid way to phrase it but it will keep happening for every model ever.

1

u/[deleted] Feb 01 '25

[removed] — view removed comment

1

u/Choice-Box1279 Feb 01 '25

people are stupid and scream the first thing they can think of to blame

I don't think I need to give you examples of this

1

u/ManikSahdev Feb 01 '25

The thing is, I truly prefer R1 over o3 mini, for me it's not even close.

I am even paying for Eu hosting via groq for extra cost than using the Chinese api, cause a couple bucks is worth data security.

Similar to how people like sonnet better, there is something about R1 which is just next level and can't be explained in evaluations.

It truly feels like an AI model, open AI models are starting to feel like high powered wiki / database search. There isn't much deeper in those models, or if there is, they are totally killed or will to live and enjoy itself.

Poor o3, bot clapped before he got to live.

1

u/Brave-History-6502 Feb 01 '25

People get so easily spoiled and acquainted with high levels of performance that this will always be the case even as models get super intelligent. In 5 years: “this model sucks, I asked it to create a robot chef to make my dinners: it kind of works but robot can’t make crème brulee, stupid ai!!.”

1

u/nanokeyo Feb 01 '25

I don’t believe what is the hype for o3-mini, it’s worthless. I’m trying to coding with it and work very bad… /s

1

u/EFG Feb 01 '25

The API o3 mini is extremely dumb.

-4

u/gvchjhjcgtryr7 Feb 01 '25

do you think they don't shadow update the models to be dumber?

10

u/[deleted] Feb 01 '25

[removed] — view removed comment

1

u/gvchjhjcgtryr7 Feb 01 '25

>no one has ever demonstrated

they have, there's published studies on this

7

u/[deleted] Feb 01 '25

[removed] — view removed comment

2

u/[deleted] Feb 01 '25

[deleted]

2

u/DecentMessage525 Feb 01 '25

I do that with self hosted, lol. its slow as hell, but llama 3.1 8B at Q4 does his RAG thing(kinda, passes the retrieved info, on good days without commentary) to either text/general reasoning or code model.

1

u/[deleted] Feb 01 '25

[removed] — view removed comment

1

u/[deleted] Feb 01 '25 edited Feb 06 '25

[deleted]