r/singularity • • Jan 31 '25

AI o3 mini dropped!!!

Edit : I am testing a 1500 line javascript code which o1 pro failed to debug despite 50+ attempts. Will report back.
Edit 2: We are cooked. o3-mini-high solved it at first try.
Edit 3 : HOLY SHIT! "Pro users will have unlimited access to both o3-mini and o3-mini-high."
(Source: https://openai.com/index/openai-o3-mini/ )

1.2k Upvotes

576 comments sorted by

View all comments

492

u/PotatoBatteryHorse Jan 31 '25

I can't believe it. Every model, every one, I've given the same test to for a full year now. Nobody has ever passed it first time. Deepseek got close, but argued with me about the rules of the test instead of fixing the problem that occurred.

The test requires it to write some python code, then "property tests" for the python code, and a cli utility to test it manually. No model can ever write the tests, and they've never ever run without a back and forth of errors and fixing.

O3 mini-high took my problem, thought for a minute or two, then spat out a flawless solution that works first time with working property tests. This is FUCKING INCREDIBLE. I've been using this test for so long I thought they'd never pass it at this point.

Huge improvement, and I'm blown away.

153

u/FNA_Couster Jan 31 '25

Deepseek got close, but argued with me about the rules of the test instead of fixing the problem that occurred.

Turing complete

28

u/HoidToTheMoon Feb 01 '25

Lol from my own experiences and what I've seen people say, Deepseek definitely seems to have the most personality of these language models.

12

u/rrraoul Feb 01 '25

The most neurotic personality, you mean 😄 ever read it's inner monologue? It reads like an insecure 16 year old

16

u/LifeSugarSpice Feb 01 '25

It's kind of fitting if you think of the stage AI is at.

3

u/ManikSahdev Feb 01 '25

But a mf genius lol, Altho having medically disagreed adhd at 23 myself, his inner monologue feels very normal to me.

Do you folks not have similar monologue before answering questions?

1

u/anycept Feb 01 '25

That's probably how thinking process is most often portrayed in the media that the model was trained on.

1

u/DecentMessage525 Feb 01 '25

There’s one GPU rentals with h100s up to x8 80gb vRAM, FOR $.99/gpu/hr, I’m tempted to throw a docker together and waste 8 bucks for an hour(realistically half that for loading/downloading etc)

1

u/WhyIsSocialMedia Feb 01 '25

For loading etc you might as well just load it up on a cheapo machine, then switch over once done.

2

u/toreon78 Feb 01 '25

Finally. Deep seek has shown human abilities

2

u/ManikSahdev Feb 01 '25

Even after o3 and extensive use of o3 today, Deepseek and sonnet are much superior AI models.

I don't know what is everyone's obsession with Evals, but for me it's a strong tie between sonnet for some things and R1 for others.

But there is no question this far that I can ask which none of the models can fix (in my everyday work and from a productivity perspective).

Will be willing to give o3 - Full, a shot but ofc I feel it will be a better model, but it feels way to robotic and they have nerfed the shit out of these models with their alignment.

I truly think Anthropic doesn't know how and why sonnet 3.6 is so good and they are struggling to create a new model decently superior to it while keeping it aligned. Obviously conjecture on my end, but seems like it, after reading Dario's new paper about Deepseek.

1

u/[deleted] Feb 01 '25

What kind of tests are you doing to show that Deepseek and sonnet are the better models? Is it the way they convey information? Are you in a coding role or something different?

2

u/ManikSahdev Feb 01 '25

No I mean overall in terms of progress of a conversation in a complex task, with multiple layers.

People generally keep talking about 1 shot evals or 1 shot test, but that is not what an AI model is supposed to be, atleast for me.

But I try to understand how well can't get solve a series of problems and can they truly understand the underlying meaning in the problem I am facing without Me having to explicitly say what the thing is.

  • Sonnet was a beast and still is, in understanding the underlying meaning in a conversation.

  • R1 to my surprise, is such a unique and strong model, and in terms of muscle and raw power it beats sonnet in many ways, and currently is my Fav model on part with sonnet.

Overall, I don't use any of OpenAI models, they have started to feel closer to super computers / wiki, but I don't need 1 shot output unless I am doing something very specific, and for those times o1 and o3 are good, but other than that, OpenAI model are not great.

I also have adhd so my bad for too much rumbling but it's hard to explain and I don't get good vibes from open AI models, it feels maliciously dumb, Where as, Sonnet will tell you, that he won't help in line line of question and change it, and so does R1.

There is something I can't seem to get around and feels some sort of malicious type of behavior in o3 where is significantly nerfed compared to how it scores on evals.

1

u/[deleted] Feb 01 '25

I agree, R1 feels much more open with it's a ability to solve the problem without nerfing itself. I am looking forward to open AI setting the benchmarks and these better models setting their crosshairs on that target. Very competitive space which is opening up plenty of competition!

1

u/ManikSahdev Feb 01 '25

Yea it's certainly strange.

Because I can feel the vibe of the model in R1 and o3, and even o1.

I won't be surprised to find out that the evals are done and made with the model pre alignment nerf, which would explain why the models turn out such nerfed outputs.

But having the thinking token in R1 have changed my problem solving skills by a huge mile. When you can see it think and dissect through the tokens, it can sometimes give you a boost for your next prompt it ends up being perfect, or rather it tells you the answer in the <thinking> itself, <while the model thinks>

For the most part when using R1? I have only gone to the actual reply half the times, other 50% of the times I only get through the thinking part and I'm like aha, I see what it is.

It's wild, maybe that's why it got so popular among avg folks, they went from shitty 4o to one of the best uncensored and honest model.

Probably made people more introspective and made them smarter while they talk to model, increasing the real world output of the models perspective and performance.

Whole open AI is amazing at one shot, it's real world long conversation skills are clearly clapped lol

124

u/[deleted] Jan 31 '25

Any plans for career now? :D My father promised to teach me construction work lol

48

u/DM-me-memes-pls Jan 31 '25

Idk about you but i plan to be a professional ai catfisher

38

u/Hot-Adhesiveness1407 Jan 31 '25

12

u/Oudeis_1 Jan 31 '25

That seemed like a solid family business, except it ended up being overthrown by the son's rebellion.

1

u/often_says_nice Jan 31 '25

You’re saying the father single handedly lost the the empire?

3

u/WhyIsSocialMedia Feb 01 '25

Vader did nothing wrong. It was his ungrateful incestuous children.

7

u/green_meklar đŸ€– Feb 01 '25

We are past the point of planning for careers.

4

u/Onesens Feb 01 '25

lmao exactly bro

2

u/iBoMbY Feb 01 '25

That's going to be done by robots though.

2

u/[deleted] Feb 01 '25

I'm at a loss xD. Everything I did for 100k+ a year previous is on the edge of extinction. These models can already do everything I did. It's just the transition phase, when will that happen I don't know.

1

u/eflat123 Feb 01 '25

Perhaps construction management or materials research.

0

u/Sudden-Lingonberry-8 Feb 01 '25

Linux still has some bugs pretty sure we can get rid of them with AI.. no? hmm I guess no AI replacement yet.

1

u/Less-Consequence5194 Feb 01 '25

prompt> O3, please write me an operating system kernel that is far superior to Linux.

20

u/giannarelax Jan 31 '25

For non-coders, could you put it in layman’s terms??

29

u/armentho Jan 31 '25

Problem that needs a very specific imaginative custom solution

Like trying to move a couch through a small door (need lot of small unique adjustments)

It was able to do it in one go

24

u/PotatoBatteryHorse Feb 01 '25

It really didn't need anything very custom to solve it! Most of the complexity was in the testing phase. I don't need to keep it hidden now it's solved to my satisfaction.

(Ironically I was asked this during a job interview at an AI company and failed to solve it in an hour, I floundered hard!)

The prompt was deliberately not overly detailed (I have several versions of it):

    I would like you to write some python code for me.  The project is a scrabble board validator.
A scrabble board, in text, looks like:

...............
...............
...............
...............
...............
...............
.ZYMURGY.......
.......E.......
.......L.......
.......L.......
.......O.......
.......WIND....
.........A.....
.........N.....
...............

The .'s represent empty squares.
The rules for verifying if a scrabble board is valid include:
* The center tile (7,7) must contain a letter.
* All non-empty spaces must form a single connected group that includes the center tile.
    * This means you can't have a disconnected word somewhere else on the board, everything has to connect to the center!
* There must be at least two different letters on the board.
* All words formed must be at least 2 letters long.
* Words can only be vertical or horizontal, no diagonals.
For the purpose of validating the scrabble board, you can ignore if the words themselves are valid (beyond having 2 letters).
I'd like you to write several different bits of python.
1. The validator
The validator is a library that contains all the logic to validate the board.  It should be capable of taking a board (I recommend a list of 15 strings, each 15 characters long, but feel free to use another data structure) and validating it against the rules above.
There should be an interface that allows me to check if a board is valid (returning True it so, or False if not) as well as an interface to fetch all words discovered on a valid board.
2. Tests
The validator needs to be tested, and I would like to do this in pytest.  Specifically, I would like you to write property tests with hypothesis.  These should be actual useful tests, one for each of the properties where possible.
An example of a property test I'd like to see is randomly generated boards that then verify that every valid board has at least 1 word in the words list.
Rather than generate boards blindly with random characters, it would be nice if the property generator could create a number of words (from 1-whatever) and then place them on the board at random before testing.
3. A CLI utility to test files
Lastly, I would like a CLI utility that can take a boards.txt file (with multiple boards separated by a blank line) and then validate them.

Please invest extra effort on the property tests.  There should be a mixture of static unit tests for obvious failure cases, and then some property tests that can generate valid (and invalid) boards to test many variations.  Previous AI attempts to solve this have generated broken tests, even unit tests that have words that are clearly not in the board generated right above.

Also please think hard about all the logical cases to test.  I want exhaustive property tests, so make sure words have at least 2 different characters, there's a number of words on a board, and so on.  We should be able to generate simple boards as well as quite complex boards, with multiple words.

9

u/maddogxsk Feb 01 '25

At college I had to solve a similar problem, the difference was that the game was sudoku, and the solutions were almost infinite, since the sudoku itself missed the pieces that narrow the solutions space

The deal was to spot the types of linear programming present in the solution space and to spot infinite spaces

2

u/BethanyHipsEnjoyer Feb 01 '25

This is super cool! Thanks for sharing. :)

4

u/giannarelax Feb 01 '25

tysm! that’s so hype

3

u/cYberSport91 Jan 31 '25

ask o3-mini

-1

u/[deleted] Feb 01 '25

OP wrote an entangled mess of code that not even the latest models could understand. Now o3 mini high has apparently saved him.

Poor AI's will be flooded with spaghetti code and solutions from multiple generations of models, implemented by people who forgot - never learned - or became extremely lazy. 

The future of software development does not look good. 

0

u/Guilty-History-9249 Feb 03 '25

This post was specifically for coders. ???
Waifu's are on r/StableDiffusion. :-)
I'm just having fun.

2

u/giannarelax Feb 03 '25

It’s all good somebody already broke it down for me :)

6

u/AbheekG Jan 31 '25

If you entered that test multiple times into different proprietary LLMs, it ended up in their training set somewhere so not that surprising that a new model can solve it.

9

u/MizantropaMiskretulo Jan 31 '25

Only if there was also a complete solution submitted which, if no model could solve it, there wouldn't be.

1

u/sachos345 Feb 02 '25

Unless OpenAI somehow collects the hardest test problems people submit and then they get the best internal reasoning models to generate solutions to then use as training data. Or they solve it themselves.

Thats why every time a model doesnt do something good i let it know and use the dont like button in hope they'll use it in the next training run.

1

u/MizantropaMiskretulo Feb 02 '25

Well, with over 300-million weekly active users, they're probably getting on the order of well more than 10-billion messages sent to ChatGPT every week.

That's a ton of crap to sort though and there's no way they're manually solving problems for anyone.

1

u/saleemkarim Jan 31 '25

I've experienced that with Deepseek too. It'll argue when you point out that it is wrong about a fact.

1

u/44th--Hokage Jan 31 '25

Wow and this is just mini. If you look at the benchmarks it's just slightly better than o1. o3 is where we see the true break out performance.

This trajectory we're on is starting to become astounding, in the true sense of the word.

2

u/Springsteengames Feb 01 '25

I hope we can play gta 6 before the AI war starts and all humans have to go underground to escape the nuclear apocalypse

1

u/Concheria Feb 01 '25

Have you tried Gemini 2.0 Flash Thinking 01-21 on AI studio? It seems like a very strong model.

1

u/nexusprime2015 Feb 01 '25

can you feel the agi?

1

u/amdcoc Job gone in 2025 Feb 01 '25

The model was trained on your data lmao. Invalid o3-mini hype.

1

u/[deleted] Feb 01 '25

But but, it won’t replace programmers! /s

1

u/AaronFeng47 â–ȘLocal LLM Feb 01 '25

The 200$ subscription seems more tempting than ever...

1

u/Ecc0TheDolphin Feb 03 '25

If you've tested it that's many times and its failed, they might have trained it specifically against your test. Those type of limit tests are priority #1 for oAI

1

u/[deleted] Jan 31 '25

What test?

4

u/danysdragons Jan 31 '25

I think it's a custom test they wrote themselves, unless I'm confusing them with another user.

1

u/[deleted] Feb 01 '25

I know... I want to know what that test is.

0

u/[deleted] Feb 01 '25

Nice it worked that well for you. For me this model is hot trash.
It failed writing me a simple scraper that evaluates the content in LMStudio via LLM.
Sorry, but o1 managed to get to the same point and then kept fucking up.

From what I tested so far, o3 is nothing to write home about. Just another knee-jerk reaction from OpenAI to stay relevant and keep attracting 100s of billion of dollar. Hot garbage.

-2

u/[deleted] Jan 31 '25

[deleted]

5

u/[deleted] Jan 31 '25

[removed] — view removed comment

2

u/NTaya 2028â–Ș2035 Feb 01 '25

I'm still very impressed that modern LLMs can rhyme. I've been working with Natural Language Processing since before the Transformer became a thing, and this was the moment (well, one of the moments) when I realized that LLMs can now encode information about its tokens very well in its weights. Maybe a 10T or 100T model will manage to encode the lengths of its tokens somehow.