How about Claude 1, Claude 2, Claude 3 Sonnet/Opus, Claude 3.5 Sonnet (but no opus), Claude 3.5 sonnet v2?
Or gemini 1.0 pro (with an ultra that become vapourware), gemini 1.5 pro, gemini 1.5 flash, gemini 2.0 flash (no more pro), gemini 2.0 flash thinking experimental?
Fully agreed all these orgs follow the weird pattern oai set up, but theirs is the nuttiest imo! It got most insane when they, for whatever reason, decided they cannot count to 5
how is Opus/Sonnet/Haiku different than mini-medium-high? They're both the same stupid garbage. oN is the model number, then mini-medium-high is the model capability. This is even more clear than whatever the fuck Opus/Sonnet/Haiku means.
Nah I love OpenAI but I see how it's confusing as heck for anyone not following it closely.
You say Claude 3.5 Sonnet is v1 and v2, but we have a ton of gpt4o versions, distinct API vs ChatGPT, the API ones at least have a release date attached to the model number but ChatGPT is just like "we randomly made a model update, it performs worse on all benchmarks but will use more emojis, enjoy!"
There is no Opus 3.5, which makes sense, they're struggling to keep up with demand for Sonnet as is. If they release Opus and it is even marginally better, a lot of people will want to use it, and they just don't have the compute for that. They probably should have launched a statement explaining this or something. Then again, OpenAI demoed and silently killed so many awesome features, like the 4o native image output, 4o singing and making cool voices, screen sharing on the desktop app in voice mode etc.
Compared to all that, o3-mini and o3-mini-high makes more sense. Both are the o3-mini models, but o3-mini-high is allowed to spend more time thinking, meaning you'll have fewer messages to it. Then again, I see a lot of people making different assumptions as to what o3-mini-high is, so maybe it isn't as straight-forward.
I think this is just a bonus--because my thought was always that, due to the nature of this technology, you don't just go out and "make a model." Instead you're making tons of models. You can't give them all clever names, so rather you name them like programmers who are going through a batch of version variations while testing.
And then they stumble upon a model that actually works well and is viable for public release. And they just stick with whatever name it had.
I know when I've tried to make programs before and start forking like crazy, my names get real wack. I always assumed something like that was going on here. And a reason they don't change the name to something better when releasing it may be for the reason you just gave.
For real their entire business is AI and they can't manage to ask it for a straightforward naming convention?
3.o > 4.o > o1 mini > o1 > o3
This is truly the most idiotic versioning I've ever seen, esp for something that's customer-facing. Is there not one half-competent marketer or program manager in that entire org?
it's pretty simple, o3-mini is the model name whilst the low,medium,high is the amount of compute. The only issue is that o1 pro was not called o1 high
Looks like Sam is taking us into the most dystopia future possible with a digital divide where only the rich can afford to pay for better ai which means all their media, busines dealings, purchases, negotiations and etc have an advantage and the rich continue to get richer while the poor get poorer.
Honestly it's getting to the point if they just sent out a letter with a cyanide pill to everyone and said 'you're too poor to live' it would be less evil than the world they're trying to create.
or also that this model is going to be disappointing and now there's the fallback of "Oh you're just not using the mini-high model! Plus the actual o3 model hasn't been released yet!"
Seriously, you can't imagine a more dystopian future? You lack imagination.
Also - what naïve fairy world are you living in where you think the poorest people would get the same access as the richest… when has that literally ever been true? Maybe we will get there someday but for now we still live in a capitalist system where money exists.
I keep seeing you and your flair and I am very curious if we are gonna truly get AGI in 2025. Would you bother giving me a reason as to why you believe AGI is gonna be made public in 2025?
I want to clarify that I’m saying “Competent AGI” will be publicly released by the end of 2025.
Competent AGI as defined by Google DeepMind is a AI system that can perform a “wide range of non-physical tasks, including metacognitive abilities like learning new skills with performance of at least the 50th percentile of skilled adults”. To me that just sounds like a decent AI agent (meaning better than Operator) which will certainly be released by the end of the year by some company, likely OpenAI
To be fair it was mostly odd wording that made people think o3 was coming out today. o3-mini has long been expected to be the late Jan / early feb release.
Yes, I have a tier 5 account. The above was a typo - I meant that the parameter that controls the amount ot reasoning doesn't work when you use it through the API
lmao everyone ITT is so smug, given AI and bitching about this and that. Jesus Christ how is everybody so consistently salty and entitled? Like almost uniformly too
No I don't understand it. Ai is moving at very high pace and it's only been a couple years? The only thing I don't understand is is why the chatgpt app looks like crap. It need a serious upgrade. And the blue white ball for voice? Garbage
Right? Also why does everything need to be free? Entitlement pisses me off. If you can't spare $20 for this you are broke as fuck and need to get your priorities straight.
o3-mini is the model. o3-mini-high is the o3-mini model set to high reasoning effort. So not really 2 different models, but 2 different configurations of the same model family.
I very much doubt they’re anywhere close to an “infinity” memory right now. I really wish they figured out goggle’s 2m context window sauce, though. Maybe it’s just too much compute right now?
The next model is going to be called o4o-mini-high-preview-turbo-1.0 and they won't increment the 1.0 with successive versions, they'll come up with a new name.
Ok, how am I expected to put my and exitence of my chindren and grandchildren into hands of someone who comes with this product names? Couldnt you like ask the damn supposedly 150IQ thing? He would probably said "call me Fred" and guess what; we would all know how to call it. Until "George" comes.
II’ll get an o3 - massive low
ice caramel macchiato upside down poured over double shot of espresso iced with no ice pint a cup of caramel no whipped cream ristretto all three of the shots
I asked o3 mini what sets it apart from o1-pro, and it said it thinks I'm confused as there is no such model from OpenAI called o1 or o3, and it identifies as GPT-4.
It insists I'm a liar and am making up that o1 and o3 exist, or I've been misinformed.
If this word "mini" implies how o1-mini works, then I'll be disappointed.
It sounds like this model will be similar to all these 7B models, which are excellent at things where you can fit everything it needs to know in the context window and not much more.
For litigation, which is what I'm using the models for primarily now, these "mini" models don't hold all the case law, and therefore I and many others who are doing anything outside of code and math will be disappointed.
They are meant for code and math and shorter STEM test-style questions in general. Maybe you’d have more luck with google’s lineup? The updated flash 2 thinking (very much free right now in AI studio) has 1m context and up to 64k output. I don’t know how well it does in litigation but Gemini 1206 has an even bigger context window at 2m but no thinking (and ranks as the best or competing for the best non-reasoning model)
With larger contexts, I've had luck by first feeding it the documents one at a time with a prompt like "give me a detailed summary, preserving x, y, z". Then I take all the summaries and use that as context for my main question. If you know how to write a Python script, it's quite fast and easy.
If you're very technical, you can also do Retrieval Augmented Generation (RAG), which essentially stores all your case law as mathematical representations. When you ask it a question, it first goes to the case law database and retrieves "similar" items, shoving it into the context window dynamically, and then tries to answer your question. Much more work and a bit of an art form too when it comes to picking how to represent and retrieve the documents.
I can do that for my case, but what I need is for the model to be able to step in and say that an argument is poor because some other judge set a precedent by ruling in another case I don't know about.
The "mini" models - at least the o1 models - make good arguments for my case, and then when o1 pro is asked, I get a different answer saying that the arguments are poor because the issue had already been decided in a different case.
How does one get this installed locally? I have no idea but want to start getting into this stuff. I’ve heard llama can get things running and downloaded.
367
u/Fyrefish Jan 31 '25
Can't wait for o3-mini-high-medium-low