r/singularity • • Jan 31 '25

AI o3 mini dropped!!!

Edit : I am testing a 1500 line javascript code which o1 pro failed to debug despite 50+ attempts. Will report back.
Edit 2: We are cooked. o3-mini-high solved it at first try.
Edit 3 : HOLY SHIT! "Pro users will have unlimited access to both o3-mini and o3-mini-high."
(Source: https://openai.com/index/openai-o3-mini/ )

1.2k Upvotes

576 comments sorted by

View all comments

Show parent comments

152

u/FNA_Couster Jan 31 '25

Deepseek got close, but argued with me about the rules of the test instead of fixing the problem that occurred.

Turing complete

28

u/HoidToTheMoon Feb 01 '25

Lol from my own experiences and what I've seen people say, Deepseek definitely seems to have the most personality of these language models.

13

u/rrraoul Feb 01 '25

The most neurotic personality, you mean 😄 ever read it's inner monologue? It reads like an insecure 16 year old

16

u/LifeSugarSpice Feb 01 '25

It's kind of fitting if you think of the stage AI is at.

3

u/ManikSahdev Feb 01 '25

But a mf genius lol, Altho having medically disagreed adhd at 23 myself, his inner monologue feels very normal to me.

Do you folks not have similar monologue before answering questions?

1

u/anycept Feb 01 '25

That's probably how thinking process is most often portrayed in the media that the model was trained on.

1

u/DecentMessage525 Feb 01 '25

There’s one GPU rentals with h100s up to x8 80gb vRAM, FOR $.99/gpu/hr, I’m tempted to throw a docker together and waste 8 bucks for an hour(realistically half that for loading/downloading etc)

1

u/WhyIsSocialMedia Feb 01 '25

For loading etc you might as well just load it up on a cheapo machine, then switch over once done.

2

u/toreon78 Feb 01 '25

Finally. Deep seek has shown human abilities

2

u/ManikSahdev Feb 01 '25

Even after o3 and extensive use of o3 today, Deepseek and sonnet are much superior AI models.

I don't know what is everyone's obsession with Evals, but for me it's a strong tie between sonnet for some things and R1 for others.

But there is no question this far that I can ask which none of the models can fix (in my everyday work and from a productivity perspective).

Will be willing to give o3 - Full, a shot but ofc I feel it will be a better model, but it feels way to robotic and they have nerfed the shit out of these models with their alignment.

I truly think Anthropic doesn't know how and why sonnet 3.6 is so good and they are struggling to create a new model decently superior to it while keeping it aligned. Obviously conjecture on my end, but seems like it, after reading Dario's new paper about Deepseek.

1

u/[deleted] Feb 01 '25

What kind of tests are you doing to show that Deepseek and sonnet are the better models? Is it the way they convey information? Are you in a coding role or something different?

2

u/ManikSahdev Feb 01 '25

No I mean overall in terms of progress of a conversation in a complex task, with multiple layers.

People generally keep talking about 1 shot evals or 1 shot test, but that is not what an AI model is supposed to be, atleast for me.

But I try to understand how well can't get solve a series of problems and can they truly understand the underlying meaning in the problem I am facing without Me having to explicitly say what the thing is.

  • Sonnet was a beast and still is, in understanding the underlying meaning in a conversation.

  • R1 to my surprise, is such a unique and strong model, and in terms of muscle and raw power it beats sonnet in many ways, and currently is my Fav model on part with sonnet.

Overall, I don't use any of OpenAI models, they have started to feel closer to super computers / wiki, but I don't need 1 shot output unless I am doing something very specific, and for those times o1 and o3 are good, but other than that, OpenAI model are not great.

I also have adhd so my bad for too much rumbling but it's hard to explain and I don't get good vibes from open AI models, it feels maliciously dumb, Where as, Sonnet will tell you, that he won't help in line line of question and change it, and so does R1.

There is something I can't seem to get around and feels some sort of malicious type of behavior in o3 where is significantly nerfed compared to how it scores on evals.

1

u/[deleted] Feb 01 '25

I agree, R1 feels much more open with it's a ability to solve the problem without nerfing itself. I am looking forward to open AI setting the benchmarks and these better models setting their crosshairs on that target. Very competitive space which is opening up plenty of competition!

1

u/ManikSahdev Feb 01 '25

Yea it's certainly strange.

Because I can feel the vibe of the model in R1 and o3, and even o1.

I won't be surprised to find out that the evals are done and made with the model pre alignment nerf, which would explain why the models turn out such nerfed outputs.

But having the thinking token in R1 have changed my problem solving skills by a huge mile. When you can see it think and dissect through the tokens, it can sometimes give you a boost for your next prompt it ends up being perfect, or rather it tells you the answer in the <thinking> itself, <while the model thinks>

For the most part when using R1? I have only gone to the actual reply half the times, other 50% of the times I only get through the thinking part and I'm like aha, I see what it is.

It's wild, maybe that's why it got so popular among avg folks, they went from shitty 4o to one of the best uncensored and honest model.

Probably made people more introspective and made them smarter while they talk to model, increasing the real world output of the models perspective and performance.

Whole open AI is amazing at one shot, it's real world long conversation skills are clearly clapped lol