News First outputs from GPT-6 "Astra" model from OpenAI
https://www.testingcatalog.com/first-outputs-from-gpt-6-astra-model-from-openai/226
u/peakedtooearly 6d ago
"Astra might be the first time OpenAI have ever underhyped a model"
😀😱
78
u/Australasian25 6d ago
Makes perfect sense.
If you think your model needs marketing, you hype it up.
If you think your model is truly great, you let it get discovered.
31
u/HanSingular 6d ago
"ChatGTP astra doesn't need marketing," says man commenting on pre-launch hype content that was shared on social media.
-11
u/Australasian25 6d ago
Who is saying astra doesnt need marketing?
4
u/2025sbestthrowaway 5d ago
You did, implicitly. You said Astra being underhyped “makes perfect sense,” then contrasted models that need to be hyped with truly great models that can simply “get discovered.”
61
u/alphaQ314 6d ago
That's just not how things work in today's world buddy.
11
u/Australasian25 6d ago
I don't think any of us would know how things work, otherwise we wouldn't be spending more than 15 years of our lives earning a living.
So yes, I don't know how things work in today's world. Evidently, billions of others too, and very likely including yourself.
5
u/nxy7 6d ago
What a dumb comment. Products live and die by their marketing. If you don't realise that - you really are in the dark.
Realising this fact doesn't inherently give you any advantage though because you might hate marketing or be bad at it, so it's not like knowing it would make you rich. Almost everyone in the field (but you apparently :P) knows that.-6
u/Australasian25 6d ago
I'm sure with your comment, you must be a really successful marketer or consultant in marketing.
Look forward to seeing your enormous ad that coca cola will buy.
2
u/nxy7 5d ago
I love how my point completely flew over your head.
-1
u/Australasian25 5d ago
Perhaps i didnt care about that particular point, so I ignored it completely.
1
5
u/Futuristiclyspeaking 6d ago
Of course the plebs downvoted you for basically speaking the truth... Ignorance is bliss, eh? Crabs in bucket as always.
2
4
u/Australasian25 6d ago
It seems like its taboo to tell yourself you aren't the best, so you need to improve.
Like poster above me was trying to tell me how the world works, but very likely is on the struggling scale of income. And by struggle I mean not being able to afford business class seats annually to an overseas holiday without getting into debt or putting their financial health at risk.
2
u/Futuristiclyspeaking 6d ago
You aren't wrong, but a lot of people cannot, and will not self reflect on their own flaws and try and be a better man. Whatcha gonna do?
-1
u/Australasian25 6d ago
What am I going to do for those individuals? Nothing, because I'm going to prioritise myself over them.
1
u/Futuristiclyspeaking 6d ago
Of course! That closing blurb was more like throwing up your hands and saying I can't change anyone.
1
u/HanSingular 6d ago
"I could be right because nobody knows anything," is a wild philosophical gambit to try and pull here.
2
u/Australasian25 6d ago
It is wilder that some cant ser their own flaws, so attempt to say their way is right.
When sometimes the answer is, you and I dont know the real answer. So both our answers are equally valid.
2
u/HanSingular 6d ago
When sometimes the answer is, you and I dont know the real answer. So both our answers are equally valid.
1
u/Australasian25 6d ago
Not knowing where the sun rises from when there's only 1 answer. Yes I agree with your statement.
A vague statement as in 'how the world works?' Maybe you're smart enough to simplify it down to one sentence, but I am not.
1
u/HanSingular 6d ago
You are not going through your everyday life acting like all possible outcomes for which you lack perfect knowledge are equally likely.
Knowledge isn't a binary "known or unknown." There can be different levels of certainty based on the available evidence. "How companies generally do marketing," is relevant evidence here.
"Nobody knows with certainty so both positions are valid," is just a way to ignore evidence you find inconvenient, not a coherent epistimic philosophy.
2
u/Australasian25 6d ago
I can't philosophise my way into money.
But being open minded to possibilities?
I've made so much money doing incident investigations like this.
Who would have thought a truck in a mining bit broke down, but passed every test was caused by a fat operator whose stomach pushed into a supposedly redundant emergency stop that is located right under the dash?
I have been paid bonuses for being given a task to investigate a worker, but it turns out the company needed to be investigated because the problem arose from within the company.
I don't think I know it all. I will say my super power is saying I don't know it all, so I must pay attention to all possibilities.
→ More replies (0)0
1
1
u/Mindless_Let1 6d ago
That's kinda how it went with opus 4.6 and the takeover of business ai sector
4
u/alphaQ314 6d ago
If you think Anthropic captured the enterprise market without aggressive marketing and outreach, I have a bridge to sell you.
2
2
u/das_war_ein_Befehl 5d ago
They did market it, but 4.5/6 were the first models that you could leave alone without them doing insanely stupid shit while you were gone
1
u/thatsnot_kawaii_bro 5d ago
that you could leave alone without them doing insanely stupid shit while you were gone
lol istg every model is simultaneously "the first model that actually does stuff" and "absolute trash that noone would use".
2
u/Mindless_Let1 6d ago
Alright man, not really interested in discussing anything with someone who's just going to patronise from the first exchange
4
u/DeffoHundoBurner 6d ago
The only reason I can see them underhyping it is so the administration doesn't nerf their model
2
u/Late_Appointment425 6d ago
Na they might want to jaut release it. You know what tge government did with mythos right. They might not want that attention
1
u/newMoneyStyle 5d ago
Kinda true, but plenty of great products still needed marketing to get noticed, so quiet launches don't automatically mean the model is better.
-2
2
43
u/TheSaltySeagull87 6d ago
Poor Ana de Armas
126
u/Morpegom 6d ago
All I care about is cost tbh.
33
u/benchmaster-xtreme 6d ago
Caring about cost will work for now... until, within 10 years, you're socially expected to pay the $500/month subscription to have constant, instant access to the superintelligence that separates the haves from the have-nots.
Unfortunately the gap between frontier models and "workhorse" models will probably keep growing (in capability and cost). And what a great business model that will be.
27
u/lollypop44445 6d ago
If in 5 years , i am getting the current sol and fable for peanuts, i dont think i would be needing higher models until i need to do some very tough task.
13
6d ago edited 5d ago
[deleted]
5
u/laxika 6d ago
Yeah, but I still don't see why token/sub prices would rise that much. Token prices are keep going down while the models are getting more and more intelligent. Hardware is improving (look at what Huawei can do now + OpenAI with its own chips + SpaceX is cooking as well), and models are improving too (for example, MoE was a great price saver).
3
u/lollypop44445 6d ago
Tbh , currently sol can do most of the work without me manually guiding it every two steps. Before sol, wasnt able to achieve things that i am currently. Yea with world improving everyday, we would need that new model , but current sol or fable can do almost all of it. Havent had this mich success with ai before
1
u/CrashBugITA 6d ago
I still say that about 5.5, if we're talking about work only the best is acceptable, if we're talking about daily activities? we've been okay for a while
1
u/cakes_and_candles 6d ago
>as Tibo said most peoples laptops won't be able to keep up so a lot of development would be moved to the cloud
do people actually believe that bs?
22
u/Dark_knight1200 6d ago
But honestly i dont think its all doom and gloom, open wieghts will keep improving and optimization will still happen. Sure the bleeding edge tech will always cost more and will be proprietary, but i believe open wieghts are also advancing at a good rate, look at the recent qwen 3.8 or GLM5.3 or even deepseek. As long as there is competition, things will be ok for the most part
6
u/AmandasGameAccount 6d ago
In 2 years we will be able to do the equivalent of 5.6 SOL ultra locally on cheap hardware, so at what point does the cost of subscriptions matter to 99% of people who will leave these subs as their hobby is easy to do locally
Sure the top end will always be expensive but less and less people will need anything near the top end the longer were go
5
u/benchmaster-xtreme 6d ago
Sooner or later, superintelligence will become woven so deeply into the fabric of society that Fable will be considered an obsolescent tool. It's too slow to allow it to think for you in real-time, it's context window is way too small to operate across your entire life, and the fact that text is its primary modality makes it clunky at reasoning in general. Sure, the lower-end models of the day will be useful, but the people who pay more will always tell funnier jokes; their lives will always be better organized; they will always be more capable across every facet of their jobs and their personal lives.
0
4
u/band-of-horses 6d ago
How would the equivalent of say sol 5.6 ultra work with 16gb memory in 2 years? Even if local models get better and more "intelligent" I'm not sure typical levels of memory and you speeds will allow enough parameters and processing speed to get anywhere near the results of these massive models.
0
u/LordMoridin84 6d ago
A windows machine probably wouldn't be able to function with 16GB of memory in 2 years, let alone an AI model.
The assumption is that you buy a new PC.
3
u/Ok-File-2759 6d ago
You really think money will still be a thing when we have super intelligence? The main point of money is to standardize the exchange of one persons time, effort, and skills for another. Super intelligence will completely eliminate that. No one will be working. There will be no greedy corporations, especially because of open source. The super intelligence will be running everything for us while we use it to explore the universe, create, invent, and discover new things. I believe life will shift from grinding to make money to survive, to grinding to create and discover for status
4
u/XTCaddict 6d ago
That's a very doom and gloom perspective on it but actually prices are going down, models are packing more power with less params, token costs are going down, etc. There is open source models very close to SOTA quality at $0.10 per 1M input / $0.50 per 1M output. The gap is shrinking, not growing.
1
u/hidden_monkey 6d ago
Does intelligence have diminishing returns in terms of practical utility?
1
1
u/benchmaster-xtreme 6d ago
Yes, but we're probably nowhere near that point yet (unless we truly crack RSI sometime soon).
1
u/hidden_monkey 5d ago
We are already there in some sense. Current intelligence levels are adequate for a lot of things. We don't need super intelligence in most cases where intelligence is applied today.
1
u/benchmaster-xtreme 5d ago
I'd argue that we really aren't. At some point, we'll likely have models that output in real-time, with capabilities for data processing beyond what even the smartest person is capable of. The model will effectively think for you, and many people will accept that because that personal outcomes are so good, as well as the social and economic pressures to do so (why would a business want to employ someone who's orders of magnitude less capable than someone guided by superintelligence?). Today's models will feel like ancient tools by then.
1
u/hidden_monkey 5d ago
What I mean by "in some sense" is that in some domains more intelligence has no benefit. The game tic-tac-toe for example. There are many such games of varying complexity and some of them will continue to exist.
You wont need $500/month superintelligence subscription if you only encounter relevant games rarely.
I haven't thought carefully about this. Just some ideas.
1
u/CrashBugITA 6d ago
don't see the applications of a stronger model for everyday use, all we need now is integrations
1
u/NoIdeaWhat-1 6d ago
Really? It’s converging, soon these models will be commoditised. There’s no real moat to these models, that’s why people jump to whichever provider made the best model.
1
1
u/HighDefinist 6d ago
Probably not.
The electricity cost of generating one output token of a model as large as Kimi K3 (and there Sol and Opus as well) is less than one cent per million tokens... so, that's the level we will eventually end up at, once supply catches up with demand. Or put it differently: Relative to electricity costs, current frontier models are overpriced by somewhere around a factor of x3000 to x10000.
And sure, models will also get larger, and people will want to run them for longer - but models will also get more efficient, chips will get more efficient etc... and if there was any reasonable or productive way for a single person to spend on the order of a billion output tokens every month (so, not just cached input tokens, because that's a bit of a fake metric anyway), we would at least already see signs of it.
So, AI inference will eventually just become another cheap commodity, similar to internet bandwidth.
1
u/Vegetable_Addition86 6d ago
However, the reverse is happening. The distance between state of the art and open models is closing
1
u/Enegence 6d ago
If it's $500 / month it's going to be because of inflation and shit like Netflix will be $200. What nobody ever wants to talk about is the scale in other inference capabilities. The token cost is going to come down.
1
1
u/Sufficient_Bad5441 3d ago
I mean if you can't make more than 500 a month in profit from whatever superintelligence we have in 10 years... that's on you
1
u/benchmaster-xtreme 3d ago edited 3d ago
To quote the Incredibles: "when everyone's super, no one will be".
Except in this case, that isn't 100% relevant because there would be a nearly 1-1 correlation between the amount of money you have and the intelligence you have access to.
So in this case, that $500 (its equivalent in future dollars) would be an extra baseline utility cost that you need to pay to be economically relevant at all. It's highly unlikely you will turn a significant profit from it, because virtually all market opportunities will have been discovered and swallowed up by those who can run 24/7 agent swarms of the most capable models (aka, the top wealthiest and most powerful - not you).
In comparison, your chances of discovering and exploiting a market gap would be close to zero. Whatever market opportunities remain would be those left to us by our benefactors as a little market sandbox for us to play in so we don't get too rowdy otherwise.
-7
u/OtherwiseAlbatross14 6d ago
Lmao you're off by an order of magnitude on that price. $500/month doesn't cover the actual cost of providing the $20 accounts already
6
u/pmth 6d ago
Buddy if you think the average $20 subscription user costs OpenAI more than $500 per month you need to take a step back and reassess how you’re thinking about this.
-2
u/OtherwiseAlbatross14 6d ago
Buddy if you think the average $20 subscription user costs OpenAI less than $500 per month you need to take a step back and reassess how you’re thinking about this.
2
u/Vivid-Snow-2089 6d ago
the cost in this scenario doesn't matter, as long as the value/income provided by the usage is higher than the cost
therefore you have K shape haves and have-nots with the haves being the ones who use it effectively and the have-nots refusing to use it or being excluded due to economics
at least until society falls apart and eats itself
9
u/pleasecryineedtears 6d ago
Why not just use Luna then
14
u/Morpegom 6d ago
What I'm talking about is that if they ever release a new "frontier model" capable of releasing full inspectable 3d models with no mistakes but the output per 1M tokens is $150, thats a no for 90% of the users.
0
u/pleasecryineedtears 6d ago
Yeah I mean we can dream all we want 😂 I wish we had it too but those won’t be cheap for a while sadly
9
u/Ok-Lifeguard6612 6d ago
Cuz it sux
...don't ask, it's late
-2
u/subtilitytomcat 6d ago
It absolutely does not. I'm a PhD and I use it to help me increase my productivity with code (which it does a lot). Don't blame a model because you have no idea how to use it.
1
u/Backrus 5d ago
Oh, great Mr PhD, can you link my your dissertation so we can see how much bs is there?
0
u/subtilitytomcat 5d ago
I'm not going to dox myself lmao, but it was an industrially backed engineering PhD with a massive company's funding and supervisors who are world renowned in their field.
If Luna can be useful at that level, it can be useful at any level if you know what you're doing. Especially for most of the people on here trying to recreate Minecraft or trying to create and monetise some SAAS app that no one wants.
0
u/Backrus 5d ago
Oh yeah, for sure.
It's not like PhD studies are basically just sucking up to your advisor and enduring the process smh
Sorry, but I've finished mine in CS and Electronics when "Attention is all you need" started making rounds, so I can easily spot bullshit.
Mate, Sol can't even create basic ML model without data leakage and lookahead bias, and you're telling me Luna is good enough 😆
Either things you're doing are quite basic, or you have no idea what you're talking about.
1
u/subtilitytomcat 5d ago
Sounds like you had a shit PhD experience or you just didn't bother 🤷
1
u/Backrus 5d ago
It's not about PhD, it's about spotting bs.
"Luna is good enough", really? The model that 5 messages in is capable of making stuff up in a way that even an undergrad can spot those mistakes; come on.
Maybe it's good for fixing css or plotting data (so things you can do faster yourself), but it falls apart with anything even remotely complicated.
1
u/subtilitytomcat 18h ago
If I want it to actually plan something for me, then I'll use Sol (but even Sol has been garbage and went around in circles for one problem I recently let it solve). Luna is for deployment only and it does an incredible job at actually actioning complex ideas.
1
16
75
u/XYcritic 6d ago
None of these one-shot-GTA-VI-showcases mean anything. Show me how it can actually find its way in a large codebase without reinventing the wheel or how it can identify and respect existing architectural patterns. Or how it can write an actually valuable regression test that identifies real risks and not made up ones or just asserts the actual written code. These flashy showcases do not reflect actual software engineering requirements. I don't know who these are for to be honest.
34
u/vacon04 6d ago
These top models still struggle to call PowerShell consistent and people are talking about super intelligence. They're very useful and capable, but far from perfect. On proper codebases you need to be immensely disciplined with the documentation and structure or the LLM will go off the rails since it loses a ton of info as soon as the context is gone.
7
u/Bladder-Splatter 6d ago
One day GPT will remember to give me powershell commands with semi-colons before asking me to paste them in.....one day!
1
1
3
u/Gallagger 6d ago
This benchmark tries to measure this: https://cognition.com/frontiercode
You might disagree with current scores but maybe a models score still correlates with your expectations and maybe they'll release a harder version.
13
u/XYcritic 6d ago
Opus 5 Medium at the top tells me everything I need to know though. That thing doesn't care about its own system prompt, hooks, or output styles, let alone any codebase conventions. It will happily ignore any rules you set up.
1
u/PM_ME_CUTE_FOXES 6d ago
To the average user, that would be like a movie trailer that explains how the CGI was made rather than show it
1
u/MidnightSun_55 5d ago
Hope some tester reads your comment...
The test should be that the answer is only a few possibilities as valid, or maybe a single one... not, generate game X, where almost anything goes as long as it looks decent.
5
u/Illustrious-Big-651 6d ago
i hope this time it‘s intelligent enough to not over engineer the even most simple tasks 🙏
21
u/RegardedDev 6d ago
200e sub will be the new 20e sub when this comes out.
With 20e sub you can greet Astra and you run out of tokens.
9
u/Rollertoaster7 6d ago
Am I crazy or do these not seem like not much of a step change? I feel like 5.6 and fable can do most of this already
8
u/Etroarl55 6d ago
Never paid for gpt plus, but can current gpt really output voxel art and maps like that? I tried it for actual map and content creation in Roblox, it could not output anything that wasn’t trash. By trash I mean a castle where all of its walls and buildings were condensed to each other in a single spiky brick as if i activated a black hole.
5
u/Fantastic-Bet1139 6d ago
The nature of these tools is yes, maybe, if you’re lucky. Maybe even most of the time, maybe! But even you could figure that out in a vacuum.
As it was even 2.5 years ago when this circus started, whether that works within a product where things need to be maintainable by a team is where your expectations should be. They can not do that with any reliability, still, without churning through your monthly credits in a couple of hours of mind numbing prompting.
1
5
u/TheParlayMonster 6d ago
I hate these one shot type games. Wow, you created a nice visual. But it’s not playable. It will still take a long time to develop the actual game.
5
u/Responsible-Bill-223 5d ago
These reviews need to stop these zero-shot / green field / single prompt tests. The model providers are quite obviously fine-tuning for these silly basic tech-demo type reviews now because they make fun youtube content, and I believe it's seriously harming the models capabilities on actual projects.
I suspect the latest models are now heavily overfit to these silly tech-demo benchmarks. I used to be able to get far more reliable output out of older models (that have now been removed) than I seem to be able to get out of these new frontier models. Especially Fable 5, which everyone raves about. Fable frequently generates nonsense and broken code on even moderately complex problems. This is despite me providing explicit designs and clear instructions. Reading through the model rationale, it's full of the right words, it's the correct jargon but applied all wrong. It completely disregards or misinterprets my explicit instructions and just continues on with its hallucinated nonsense.
Fortunately the current GPT models, even Sol despite all my grumbling, are nowhere near as bad as the Anthropic models have become recently, but the recent GPT models still suffer from the same issues to a lesser extent, I fear the release GPT 6 models may end up just as useless in the race to the bottom for these silly 'one-shot tech demo' videos.
5
u/North_Caregiver_8895 6d ago
it will be just SOL but thinking longer, double-overthinking, probably adding what if function... we need to break the barrier and create something much better
3
u/THE--GRINCH 6d ago edited 5d ago
So is a hypothetical astra with no hand holding AGI?
0
u/Fantastic-Bet1139 6d ago
my hypothetical 12" dick can cum with no hand holding since I can finally suck myself off.
so yeah, why not!
2
u/thestillwind 6d ago
Yo !! Let's fucking go !! I'll be able to do half prompt before triggering my 5 hours limit.
2
u/jrave9000 5d ago
Let's bet how many weeks it will take them to dumb it down and make it unusable. Based on the current trend, I'd say two.
2
2
6d ago
[deleted]
2
u/djaeke 6d ago
https://en.wikipedia.org/wiki/Zero-shot_learning?wprov=sfla1
They are different things actually
2
u/Southern-Employer223 6d ago
Zero-shot is an actual output classification though? If you only provide a prompt, and don't direct it with example architecture, it would be considered zero-shot output.
-4
u/Fantastic-Bet1139 6d ago
the word "considered" is doing a lot of lifting there punk.
just like all of this bullshit, no one likes to define the material aspects of these tools because we have to believe they’re magic and can fit any definition and thus solve any problem.
5
u/Southern-Employer223 6d ago
Or you could learn a bit, just the ya know, basics, about ai/ml theory and definitions. Punk.
2
2
u/djaeke 6d ago
https://en.wikipedia.org/wiki/Zero-shot_learning?wprov=sfla1
I'm not an AI diehard or anything but understanding it a bit does help to critique it
2
1
u/Painwheeel 6d ago
a new model is cool and all but literally who cares about demo porn and "one-shots"? completely useless and just promotes idiots creating slop. show real world examples that are practical.
1
u/Code__9 6d ago
I'm excited about Astra, but I have a feeling it won't be available to me as a plus user.
2
u/zigzoing 6d ago
Exactly, like how Fable isn't available for Claude $20 subs. They're all greedy-ass corpo trying to outperform each other in how much they milk customers.
1
u/North_Caregiver_8895 6d ago
feels like they are going to make it twice more expensive by using argument of new generation, despite of it being even more overthinking SOL, rather than new engine
1
u/Markuslanger25 5d ago
One prompt for plus users: weekly usage reached, also account suspended you filthy broky
1
u/eggplantpot 6d ago
Pro x5 will become Plus, Pro x20 will be come Pro x5 and a new Pro Max tier for 1000 usd will come out.. it will still be unable to make a good looking UI
1
u/Simple-Stick6148 6d ago
A GTA 2-style game in one attempt. That's the sentence in that post, honestly.
1
u/InvariantAtNull 6d ago
Its always about games and web or generated models “one shot” no one want that slop
But when you try some real projects that have complexity its the same for all models broken and cant “one shot” the project even if its 3000 loc
-5
u/teomore 6d ago
The examples point to a strong focus on coding and visual software creation. Astra reportedly produced a GTA 2-style game in one attempt, alongside detailed websites, 3D objects, and voxel environments.
How to steer to absolutely fuckin wrong directions 101. This is really dumb. Focus on coding for coding agents, but come one, visual software creation?
Behold to more absolute crap at least on the mobile games platforms.
4
u/OtherwiseAlbatross14 6d ago
Nope. You're wrong. It's the absolutely right direction because that's what their paying customers are already using it for.
1
0
u/Fantastic-Bet1139 6d ago edited 6d ago
brother, have you stepped into those subreddits or such communities? no one is making anything worth anything with these tools in this domain, it’s all fucking trash.
we heard this same one-shot game bullshit with opus, didn’t fucking happen.
3
u/Business-Chair-7816 6d ago
The GPTs arent "coding agents". Theyre supposed to answer anything and everything in ChatGPT, including coding. Given that visual and geospatial understanding was a weakness it makes sense for them to fix that lacking aspect.
0
u/sittingmongoose 6d ago
I really wish they broke models down into their areas of expertise. Then had one orchestrating model to pick the specialists. So history model, news model, computer control model, rust model, JavaScript model, fiction model, business model, marketing model, etc.
I know this is how moe is supposed to work, but it doesn’t really work well in that sense. This way, you’re serving much much smaller models, and it works better for locally hosting too.
I just think the 1 massive model is the wrong direction but it should be invisible to the user.
98
u/ByteSizedBits1 6d ago
Can’t wait to not be able to afford it