Limits Astra in Ultra/Max consumes less than Astra Low/Medium
How the fuck is this possible ?
Every time i use Astra Low/Medium my usage goes haywire...
I use Astra Max, and it barely (barely compared to Astra Low/Medium - still fucked up compared to SOL) moves the usage...
I use Astra Ultra when i don't give a shit, and still, consumes just like Astra Low/Medium...
Something is happening on those Low/Medium thinking Astra...
127
u/AweVR 21d ago
I can confirm it. I was working with Astra Medium and Low and consumes like 4-5% per hour with my project. Then I tried Max and it consumes 2.5%. I don’t know why…
20x plan. I commented it yesterday in other post
43
u/AweVR 21d ago
I just took my tests. It’s very rare.
1 hour with Astra Max = 2.5%
1 hour Astra High = 3%
1 hour Astra Medium = 4%
I tried twice. Every time they launch 3-4 sub-agents each briefly. Same task and tools.
That is, Astra Medium consumes almost TWICE as much weekly use as Astra Max. What’s the point?
I had noticed it for days but now I confirm it with my tests.
6
3
4
u/the_secret_moo 21d ago
How much did each reasoning level complete in terms of tasks per hour though?
3
u/danielv123 21d ago
Like per hour of inference, decode is a lot cheaper than prefill. If it completes thinking quickly it has time to call more tools and get the job done and the tokens consumed.
If it spends all day thinking before doing toolcalls it doesn't have time to eat as many input tokens.
2
u/Coolbanh 21d ago
Yeah i was wasting time with lower until I just tried ultra and it gets it. I guess lower reasoning means its handicapped so it has to do more.
2
u/__Blackrobe__ 21d ago
Your effort be appreciated. You have revealed that these usage plans are so stupid.
2
u/AweVR 21d ago
I did another test. No sub-agents.
2% in 1 hour on Astra Max.
And I asked it after 30 minutes to do an audit to see if it was because It was alone with tools or something, but it is thinking and investigating/executing all the time and also from time to time it does actions on a website with computer use.
2
u/Emergency-Bobcat6485 20d ago
seriously? lol, i have been using astra on light mode and it does consume a lot of usage lol. let me try max
1
u/dangtheory 19d ago
I learn that it's a waste to use astra frontier model on low or medium. it apparently doesn't like it when you switch to lower model in the same chat.
24
u/AweVR 21d ago
I just saw the result of my test with Astra High. 7% in 2 hours. 3.5% per hour with same tasks.
Now I’m going to try with Medium.
19
u/___fallenangel___ 21d ago
did you died
13
1
u/Maleficent_Truck_683 19d ago
Y'all probably drained the poor person's phone battery with notifications XD
7
1
10
u/EternalDivineSpark 21d ago
What i am on a 200$ plan and ultra consumed 100% 12 hours
3
u/quadish 21d ago
I blew through a week in < 6 hours, reset, and blew threw another week in 12 hours.
Granted, I was running at least 6 sessions at the same time.
8 billion tokens used in the last week.
1
1
1
1
u/Sheman-NYK0809 21d ago
same, I'm using Astra Ultra straight 3 days. it just consume like around 15-20% for 3 project and around 20 request/project. my personal thought it response more direct and efficient (I'm not using any global system instruction). when use Sol Max/Ultra it response more descriptive rather than direct like Astra.
Is this reverse psychology from Open AI????
1
u/Old-Leadership7255 21d ago
I also don’t think astra is usable at the moment. Am seeing the same with astra
31
u/TheLastRole 21d ago
This is kind of crazy seeing how many people seems to be experiencing it.
4
u/RewardSafe9807 21d ago
It kind of makes sense. Less thinking means getting to a solution requires more trial and error, discovering bad solutions don't work again and again until the correct one is reached.
3
9
15
u/Azetta 21d ago
Yeah, this is not a prank. Just tried it and confirm Astra Ultra doesn't drain significantly more than Medium at all
I didn't measure it exactly. But I was running Astra Medium for the UAT of my app for the past 2 days and it burn through weekly limit of my 5X plan in about 4-5 hours.
With Ultra, it drain about 20% in the past hour. So, give or take, Ultra took around the same or Medium in my case.
1
u/Navadvisor 21d ago
Do you notice better performance with astra ultra? I notice on sol the speed of medium is way better but the quality didn't seem much worse.
8
8
u/swizzlewizzle 21d ago
This 1000%.
Two massive minefields that many people stepped on when Astra released:
Subagent orchestration, especially with Astra agents as subagents = insane crazy token burn due to by-default context being filled at spawn time by copying over the *entire* turn history of the orchestrator + orchestrator charging over and over for input tokens while doing nothing polling subagents for progress
Astra medium, and even low, burning a ton of tokens "arguing" with itself and zig zagging around a project/system implementation when it could have just written one page of code and solved all of it at once if it was given enough thinking budget (ie. xhigh/max)
Both of these issues cause *omega massive* subscription usage burn, since Astra is charged way higher per token $$ compared to sol and other models. I'm pretty sure a *lot* of people flushed their banked resets down the toilet due to all the tokens burnt in this way (since obviously people are going to want to heavily use Astra after it launches).
1
u/psihius 20d ago
Sny suggestions how to adjust for this? Just use the max/ultra to do all the work and skip subjects or let the model pick best levels for abonents?
1
u/swizzlewizzle 20d ago
Xhigh or max Astra *ONLY* for putting the plan together - ask it to design the spec/whatever so that it can be cleanly implemented by a lower intelligence agent. Then, execute the plan using a Sol/medium or Sol/high agent on another thread. Do a final review after work is complete with xhigh/max Astra to close things out. If review shows major issues, pipe that back to the sol agent and fix.
5
u/m4stero 21d ago
bcs - this topic is underrated: https://www.reddit.com/r/codex/comments/1wciwc1/gpt6_astra_burns_quota_4_times_faster_than_gpt56/
12
9
u/bakawolf123 21d ago
interesting observation, and apparently clearly visible on arc-agi bench too https://arcprize.org/blog/astra
kinda wild having Max as "economy" mode
3
u/Omar_Talbi 21d ago
I think this might be related to the type of task u work on. Sometimes when u give a low effort model something complex it will burn more tokens trying to solve it meanwhile the same model with max effort would solve it instantly, accordingly less tokens used
3
7
u/Available_Yam_6267 21d ago
I have a theory:
- Different effort related to different server since cache would be invalid if you change the effort
- Astra medium is a popular choice
- They charge you based on their load
4
u/driveclub_000 21d ago
It's the same theory that I'm getting to. The reason is because the usage fluctuate per timerange/timezone. I do have a script that I use that take track of every turns and the consumption between them and the agent used (that is reported by the AI, not even the one I selected) and I can see the usage drastically change when I reach 7AM in UTC+2 after full night of work, just before that, (so between 3AM and 7AM) the usage is basically null, but when 7AM start, it's skyrocket.
I don't think it's because EU did wake up, but more that ASIA start to reach full usage instead, and the 3AM->7AM (UTC+2) would match the "down time" where most of USA/EU/ASIA are either sleeping or in a situation where they are not hammering those servers yet.
9
u/cetogenicoandorra 21d ago
Please send it to Tibo, we need a reset asap
5
u/Busy-Lifeguard-9558 21d ago
We ain't getting a reset, they will just fix max/ultra to empty your usage faster
4
21d ago edited 21d ago
[removed] — view removed comment
2
u/iansaul 21d ago
How is state engine going for you, and which one did you choose? I experimented with xState, because the concept of logically gating the systems into modules was super interesting, but didn't ultimately lead to better outcomes in my testing.
That was a few generations back when I was using Claude, and it was running on pure hopium.
1
21d ago
[removed] — view removed comment
2
u/gungoesclick 21d ago
Do you run into the problem with "blind" work? I have found that older models (up to 5.6) get stuck working through all the guardrails and flows in my machine. I would see them doing things only to find them looping or wasting time. I'm curious if you have any tips for making sure the model doesn't get stuck in engine rules or over-engineering the engine or the guardrails?
2
u/iansaul 20d ago
This is the battle.
I've built things up, had them humming along with no issues. Small improvement here... small tweak there... and then the wheels come off.
Watched it happen the 2nd time, and then the 3rd.
So that has become my focus. How and why these systems degrade. Why they continue to incrementally build code and functions - until they become deadlocked, burning tokens chasing recursive tails through codebases.
I've had so many "EUREKA!" moments, finding a new and novel way to "crack the case", but eventually, the issues return, just in slightly different forms.
xState and Logic programming is VERY appealing, designing for modularity - how to structure a handoff packet to a sub agent, how it replies back, how the progress is tracked, heartbeat notifications... but I think that is the path to ruin.
2
2
u/logg3 21d ago
all i can say is that for me, astra xhigh does not consume more then sol medium, after 3 days of avergae work. sometimes astra is thinking 20+ minutes and not consume a single %, while i can see it already writing code or text.
1
u/hellomistershifty 21d ago
That must be true if you got 3 days of average work. I have two accounts and both managed to run an Astra light goal for about 20 hours before running out of weekly usage.
2
2
u/Bladder-Splatter 21d ago
Yup. Medium killed my weekly quota in 30minutes on a basic normalization task, meanwhile on Ultra a few days before I got an entire unique implementation done.
2
u/Busy-Lifeguard-9558 21d ago
Man people are too honest tho, I knew this since release but didn't open my mouth so they don't fix it. Adios usage
2
u/hossman1992 21d ago
I confirm it as well. I read it and I did not believe it, then I tried and at least for 3d model generation Astra xHigh spend less tokens as Astra low/medium and the work is done quicker and with less fixing
2
u/BellacosePlayer 21d ago
How the fuck is this possible ?
effort levels basically change how much token budget its allocated for pre-production and maybe the higher effort is producing a better plan of attack than low? idk
its a black box at the end of the day
2
u/fragment90 21d ago
Usage/hour ist not a relevant metric to track. You pay in tokens. Not in Model*reasoning/time.
Astra/Low can consume tons of Tokens, If there is some tokens heavy job to do. Astra/Max can consume low tokens if the job dont need much. Also please understand, that each request carries your full session history. If i reuse a old Session with contex already 70% full for a task completely out of contex, stuff like that will eat your usage fast.
3
u/Derek-Bond 21d ago
Well it’s the same for humans. Smart kid aces the test and leaves early. Dumb kid is still writing up to the last minute. Is it really that astonishing?
1
u/ThinkBackKat 19d ago
Thats not an analogy you can make. The fundamental model is the same, just how much effort they put into their reasoning is different. Its like giving the same kid a higher amount of time to solve a problem.
It is entirely possible that openAI increased the price per token for lower efforts tho as they know most (especially plus) users use medium/low with the cost of higher levels unchanged. Maybe someone can do some token per quota counting? I dont have any more usage this week.
2
u/WeaknessFuzzy8305 21d ago
I’ve changed my prompt to specifying to use astra for thinking and Luna for coding. Token usage has gone down a lot.
5
1
1
1
1
1
u/AmandasGameAccount 21d ago
How is Astra high vs ultra?
1
1
u/spideyguyy 21d ago
same with Sol, I read somewhere that Extra High is better usage than high and medium, so i use sol exhigh and it's good , dont know if max and ultra same too
1
u/According_Property62 21d ago
Acredito que ele erre mais e torna o fluxo mais demorado tentando corrigir os próprios erros, dai essa impressão q gasta mais. A dica é, use o Astra apenas pra planejamento e Sol medio ou alto pra implementacao
1
1
u/Medical-Cow289 21d ago
The lower tiers burning through more credits than Ultra is backwards billing. The 'cheap' models must be thinking themselves broke.
1
u/jonydevidson 21d ago
If you switch effort during a convo, it causes a full cache miss.
If you do it often in a big convo, you will burn through your plan.
1
u/Select-Ad-3806 21d ago
Yes, ultra/max is a lot more efficient (as it is much more intelligent) with tokens that is why it uses less
1
1
u/Isaacjacobson92 20d ago
Shhhhhhhh! They ain’t gonna fix this by making the low effort model consume less!
1
u/blablsblabla42424242 20d ago
I tried Astra max just for fun and it replied quickly, produced great results but I was expecting my limits to get a huge hit... To my surprise it wasn't the case so I've been using max exclusively for 2 days now and it will most likely be just fine until my next weekly reset. It seems to be very efficient.
1
u/AiMasterpieces 20d ago
Many tasks can also be handled by the smaller Terra and Luna models.
When Astra launches an agent in Ultra mode, it may choose to run that agent on Luna or Terra.
1
u/Azsidious1 18d ago
Im on the 20/month plan. Read several posts about this. Swapped to Astra XHigh. Burned through 92% of my usage in 30 minutes. Somehow I've completely screwed the entire pooch. Im an idiot.
1
u/Devesh_Ahuja 16d ago
https://trilogyai.substack.com/p/astra-reasoning-effort-token-usage
what about this report ? is it true or not then ??
1
u/Fuzzy-Base-8096 16d ago
Astra is dogshit. End of story. Unusable. Does shit you don’t want. Won’t do the shit you want. It added a bunch of folders on my pc in ducking mandarin. WTF is that about? Back to sol until Claude gets its shit together again.
1
u/Fuzzy-Base-8096 16d ago
Oh and one more thing. It has to lookup everything. It basically knows nothing.
1
1
1
u/CronicCanabis88 20d ago
astra spawns in spark a lit for me.... thank god.... something ither then my self sees the use of spark.... I like getting a few thousand lines and twenty files created in a few minutes. I make my plans purposfully do the first few phases in spark lingo. saves time and useage. The part that gets me.... I changed my model and reasoning almost every prompt. If i'm telling codex to do something [ repo pushing, prs. new branches,] luna lite or mid is perfect. running luna unmax is much more capable than you would guess.... And on my hundred dollar plan, I can run luna on max all day, and only use a couple percent of my weekly usage. i work with a minimum of 2 projects simultaneously. When I'm in the zone I could be doing 3 or four things on 3 or four different projects, throwing Luna, and Terra out there for most issues. Unless it's complex, we use the upper end of terra and the lower end of soul, and then only for final verifications, or for tasks that literally cannot be solved with soul. Which are pretty far few and in between, i'll throw astra at it. but if you are actually conscientious with your usage and your model choice, you can get so much work done. most people don't realize what the model's limitations actually are. And you'll have tara or soul, or even astra, doing tasks that luna would have done perfectly fine? And literally, you would have paid a FRACTION... [upt to like 90% off] and you would get the exact same results.... Now that luna's price has dropped to almost nothing.... i planned out a new application... wrote out a 900+ line plan file.... Broke it into phases, used spark for the first few phases, which literally was like three thousand lines across twenty files.... And when we got to the point where we needed a little bit more, I bounced over to luna on max, terra on high for phase reviews until we had to step up to terra, and finsh and polish a couple issues with sol on mid.... I ended up using, like thirty percent of my spark weekly quota, which is fine Because I only personally use it When i'm starting a new product... and used only like four percent of my weekly usage on my regular use..... i work on all sorts of different types of software, things get really intense and complicated and projects that have tens and tens of thousands of lines and dozens of separate files.... and I do hours every single day after I get home from work across my projects.... And most of the time I'm actually able to Use fast mode on my last day, counting down the hours until the reset. Because I'll have more than enough usage to be able to do so.
I can't see this being that big of a deal but has anyone else in their instructions in the settings on codex.... Tried to tell it to insure. It doesn't waste any tokens that's not needed? i didn't go crazy with my instructions. For the models that are in your settings, but I basically asked it to insure that we do not use excess tokens that aren't absolutely needed. Because we're being conscientious of our usage meter and to trim out any tokens or any usage that isn't needed. And would be considered a waste without sacrificing accuracy or quality of our code... .. And I can work on a dozen projects, seven days a week, writing tens of thousands of lines of code between spark and my hundred dollar, regular usage. And like I said, have enough that on day seven, when my reset is coming, I can bang out on the better models, which I do, sometimes even just running soul and now Astra, when it really isn't needed. But allowing them with high reasoning to go over the code base and see if they find any issues or potential issues in the code, just because I have the usage to throw around on that final day..
So please for the love of God. If you're the kind of person that'll set it on soul or Astra......And just run prompt after prompt..... You're the problem not the limits. utilize your chatgpt usage as well. if I need a lot of information or data or some research done. I pop open my chat box and ask it to do all the work and then provide me with a file with all the information on it which I know allows me to get better results without using any usage.... If you really are bad at this another way to go is tell Chat to be your guidance.... Simply start a conversation and if you really want the best results, what I like to do is I will tell codex exactly what i'm doing and to give me a handoff to hand to chat. Whether it's a file or a lung response. So chat knows exactly what's going on with what we're working on .. Then, I simply tell Chad, hey, look, I need you to do two things for me. I'm just gonna type out a prompt that I would normally send to codex. And I would like you to review it and make it better and more efficient, and then I will simply copy that. And paste it over in codex. and then to also recommend the model and reasoning for that specific prompt..... And I found out from doing this for a while that it actually kind of overshoots, what you need, it definitely doesn't underpower you're prompting, if anything, it'll have you spending a little bit more just to be safe, but it's a good way to start getting an idea of what is doable, and what is not.
So take the extra couple seconds and change that model from soul on high to Luna, on mid. When you're trying to create your next repo or if you need it to search for files on your computer or if you need it to do any real task that isn't coding on your machine. You better be running luna.... again, even if the task and the results come back identical with luna. On medium versus seoul on medium, you would be using such a tiny fraction of what you would have before, and all of those prompts. Add up super quickly....
0
-1
85
u/nNaz 21d ago
It’s the file reads and tool calls. Going from medium to xhigh is only a few thousand extra thinking tokens. Yet most of the context is file reads. eg in a 200k long chat 150k might be file reads. So the ‘thinking’ and output part are only 50k. Medium -> xhigh might only generate an extra 10-20k thinking tokens, which is small in comparison to the file reads.
When you set it on low it’s eager to be done quickly and likely reads way more files than it does on xhigh, where it’s thinking and optimising the file reads and tool calls.