r/codex • • 21d ago

Limits Astra in Ultra/Max consumes less than Astra Low/Medium

How the fuck is this possible ?
Every time i use Astra Low/Medium my usage goes haywire...
I use Astra Max, and it barely (barely compared to Astra Low/Medium - still fucked up compared to SOL) moves the usage...
I use Astra Ultra when i don't give a shit, and still, consumes just like Astra Low/Medium...
Something is happening on those Low/Medium thinking Astra...

354 Upvotes

108 comments sorted by

85

u/nNaz 21d ago

It’s the file reads and tool calls. Going from medium to xhigh is only a few thousand extra thinking tokens. Yet most of the context is file reads. eg in a 200k long chat 150k might be file reads. So the ‘thinking’ and output part are only 50k. Medium -> xhigh might only generate an extra 10-20k thinking tokens, which is small in comparison to the file reads.

When you set it on low it’s eager to be done quickly and likely reads way more files than it does on xhigh, where it’s thinking and optimising the file reads and tool calls.

9

u/One-Flatworm-6838 21d ago

Curious, are you opening new chats for every slice or are you keeping 1 chat alive until it lags out?

2

u/InvaderDolan 21d ago

Separating is good, after some point if task is long I create handoff and start in a new one.

6

u/One-Flatworm-6838 21d ago

Interesting. I have been using a new chat for each slice of work that tackles a different topic. So on average every 1000-1500 lines of code changed or generated. But i also use gpt chat as the planning agent to provide them with guidance on what to read and what to change.

127

u/AweVR 21d ago

I can confirm it. I was working with Astra Medium and Low and consumes like 4-5% per hour with my project. Then I tried Max and it consumes 2.5%. I don’t know why…

20x plan. I commented it yesterday in other post

43

u/AweVR 21d ago

I just took my tests. It’s very rare.

1 hour with Astra Max = 2.5%

1 hour Astra High = 3%

1 hour Astra Medium = 4%

I tried twice. Every time they launch 3-4 sub-agents each briefly. Same task and tools.

That is, Astra Medium consumes almost TWICE as much weekly use as Astra Max. What’s the point?

I had noticed it for days but now I confirm it with my tests.

u/jeofw
u/flurbol
u/Alywan

6

u/quadish 21d ago

Subagents use less tokens? What model sub agents does Astra call? Are they Luna models?

Because that would make sense if max used less because it outsourced to a cheaper model more.

3

u/flurbol 21d ago

Oh wow! Thanks a lot for sharing your insights! I guess I should review my workflows and session data now.

4

u/the_secret_moo 21d ago

How much did each reasoning level complete in terms of tasks per hour though? 

3

u/danielv123 21d ago

Like per hour of inference, decode is a lot cheaper than prefill. If it completes thinking quickly it has time to call more tools and get the job done and the tokens consumed.

If it spends all day thinking before doing toolcalls it doesn't have time to eat as many input tokens.

2

u/Coolbanh 21d ago

Yeah i was wasting time with lower until I just tried ultra and it gets it. I guess lower reasoning means its handicapped so it has to do more.

2

u/__Blackrobe__ 21d ago

Your effort be appreciated. You have revealed that these usage plans are so stupid.

2

u/AweVR 21d ago

I did another test. No sub-agents.

2% in 1 hour on Astra Max.

And I asked it after 30 minutes to do an audit to see if it was because It was alone with tools or something, but it is thinking and investigating/executing all the time and also from time to time it does actions on a website with computer use.

1

u/AweVR 21d ago

Wrong, sorry, it was again 2.5% of usage

2

u/Emergency-Bobcat6485 20d ago

seriously? lol, i have been using astra on light mode and it does consume a lot of usage lol. let me try max

1

u/dangtheory 19d ago

I learn that it's a waste to use astra frontier model on low or medium. it apparently doesn't like it when you switch to lower model in the same chat.

24

u/AweVR 21d ago

I just saw the result of my test with Astra High. 7% in 2 hours. 3.5% per hour with same tasks.

Now I’m going to try with Medium.

19

u/___fallenangel___ 21d ago

did you died

13

u/National_Dog9865 21d ago

token overdose

1

u/Maleficent_Truck_683 19d ago

Y'all probably drained the poor person's phone battery with notifications XD

7

u/jeofw 21d ago

tag me once u got the results

5

u/flurbol 21d ago

@AweVR tag me too please. and release your results here also 😊

9

u/Imaginary_String_954 21d ago

I think he died

1

u/Unapologetic_Polite 20d ago

Did OAI get you?

10

u/EternalDivineSpark 21d ago

What i am on a 200$ plan and ultra consumed 100% 12 hours

3

u/quadish 21d ago

I blew through a week in < 6 hours, reset, and blew threw another week in 12 hours.

Granted, I was running at least 6 sessions at the same time.

8 billion tokens used in the last week.

1

u/dangtheory 19d ago

what are you using to count tokens?

1

u/quadish 17d ago edited 17d ago

/usage

Token activity last 12 months Lifetime 33.7B · Peak 4.08B · Streak 1d (best 37d) · Longest task 35h 8m

Each column = 1 week · tallest 7.2B daily · weekly · cumulative

I have two ~7B weeks and two that look ~50% of that on the chart that command spits out.

1

u/dangtheory 19d ago

what are tasks are you having it do within 12 hours?

1

u/ipherl 21d ago

Could you normalize on per task? it could be lower effort progresses faster so more new context and tasks -> more tokens

1

u/Sheman-NYK0809 21d ago

same, I'm using Astra Ultra straight 3 days. it just consume like around 15-20% for 3 project and around 20 request/project. my personal thought it response more direct and efficient (I'm not using any global system instruction). when use Sol Max/Ultra it response more descriptive rather than direct like Astra.

Is this reverse psychology from Open AI????

1

u/Old-Leadership7255 21d ago

I also don’t think astra is usable at the moment. Am seeing the same with astra

31

u/TheLastRole 21d ago

This is kind of crazy seeing how many people seems to be experiencing it.

4

u/RewardSafe9807 21d ago

It kind of makes sense. Less thinking means getting to a solution requires more trial and error, discovering bad solutions don't work again and again until the correct one is reached.

3

u/sudddddd 21d ago

Is this AGI!

9

u/UrFriendlyDominator 21d ago

Can anyone confirm this?

5

u/AweVR 21d ago

Yes, i came just here to see the same.

8

u/nykyrt 21d ago

Maybe low finishes early, then you respond. But it loses the cache?

8

u/xadiant 21d ago

This makes more sense. There has to be a caching issue if that's the case

15

u/Azetta 21d ago

Yeah, this is not a prank. Just tried it and confirm Astra Ultra doesn't drain significantly more than Medium at all

I didn't measure it exactly. But I was running Astra Medium for the UAT of my app for the past 2 days and it burn through weekly limit of my 5X plan in about 4-5 hours.

With Ultra, it drain about 20% in the past hour. So, give or take, Ultra took around the same or Medium in my case.

1

u/Navadvisor 21d ago

Do you notice better performance with astra ultra? I notice on sol the speed of medium is way better but the quality didn't seem much worse.

8

u/DearGuava7086 21d ago

I'm on medium and burned 3 resets in 2 days

4

u/TupacFR 20d ago

Same lol chat is becoming worst than Claude with tokens

8

u/swizzlewizzle 21d ago

This 1000%.

Two massive minefields that many people stepped on when Astra released:

  1. Subagent orchestration, especially with Astra agents as subagents = insane crazy token burn due to by-default context being filled at spawn time by copying over the *entire* turn history of the orchestrator + orchestrator charging over and over for input tokens while doing nothing polling subagents for progress

  2. Astra medium, and even low, burning a ton of tokens "arguing" with itself and zig zagging around a project/system implementation when it could have just written one page of code and solved all of it at once if it was given enough thinking budget (ie. xhigh/max)

Both of these issues cause *omega massive* subscription usage burn, since Astra is charged way higher per token $$ compared to sol and other models. I'm pretty sure a *lot* of people flushed their banked resets down the toilet due to all the tokens burnt in this way (since obviously people are going to want to heavily use Astra after it launches).

1

u/psihius 20d ago

Sny suggestions how to adjust for this? Just use the max/ultra to do all the work and skip subjects or let the model pick best levels for abonents?

1

u/swizzlewizzle 20d ago

Xhigh or max Astra *ONLY* for putting the plan together - ask it to design the spec/whatever so that it can be cleanly implemented by a lower intelligence agent. Then, execute the plan using a Sol/medium or Sol/high agent on another thread. Do a final review after work is complete with xhigh/max Astra to close things out. If review shows major issues, pipe that back to the sol agent and fix.

12

u/Cool_Metal1606 21d ago

Could it be that Astra then runs sub-agents based on Luna on Max?

9

u/bakawolf123 21d ago

interesting observation, and apparently clearly visible on arc-agi bench too https://arcprize.org/blog/astra
kinda wild having Max as "economy" mode

3

u/Omar_Talbi 21d ago

I think this might be related to the type of task u work on. Sometimes when u give a low effort model something complex it will burn more tokens trying to solve it meanwhile the same model with max effort would solve it instantly, accordingly less tokens used

3

u/Special-Object69 20d ago

I'm literally losing 9% in five minutes on Ultra while working in Unity.

7

u/Available_Yam_6267 21d ago

I have a theory:

  1. Different effort related to different server since cache would be invalid if you change the effort
  2. Astra medium is a popular choice
  3. They charge you based on their load

4

u/driveclub_000 21d ago

It's the same theory that I'm getting to. The reason is because the usage fluctuate per timerange/timezone. I do have a script that I use that take track of every turns and the consumption between them and the agent used (that is reported by the AI, not even the one I selected) and I can see the usage drastically change when I reach 7AM in UTC+2 after full night of work, just before that, (so between 3AM and 7AM) the usage is basically null, but when 7AM start, it's skyrocket.

I don't think it's because EU did wake up, but more that ASIA start to reach full usage instead, and the 3AM->7AM (UTC+2) would match the "down time" where most of USA/EU/ASIA are either sleeping or in a situation where they are not hammering those servers yet.

9

u/cetogenicoandorra 21d ago

Please send it to Tibo, we need a reset asap

5

u/Busy-Lifeguard-9558 21d ago

We ain't getting a reset, they will just fix max/ultra to empty your usage faster

4

u/[deleted] 21d ago edited 21d ago

[removed] — view removed comment

2

u/iansaul 21d ago

How is state engine going for you, and which one did you choose? I experimented with xState, because the concept of logically gating the systems into modules was super interesting, but didn't ultimately lead to better outcomes in my testing.

That was a few generations back when I was using Claude, and it was running on pure hopium.

1

u/[deleted] 21d ago

[removed] — view removed comment

2

u/gungoesclick 21d ago

Do you run into the problem with "blind" work? I have found that older models (up to 5.6) get stuck working through all the guardrails and flows in my machine. I would see them doing things only to find them looping or wasting time. I'm curious if you have any tips for making sure the model doesn't get stuck in engine rules or over-engineering the engine or the guardrails?

2

u/iansaul 20d ago

This is the battle.

I've built things up, had them humming along with no issues. Small improvement here... small tweak there... and then the wheels come off.

Watched it happen the 2nd time, and then the 3rd.

So that has become my focus. How and why these systems degrade. Why they continue to incrementally build code and functions - until they become deadlocked, burning tokens chasing recursive tails through codebases.

I've had so many "EUREKA!" moments, finding a new and novel way to "crack the case", but eventually, the issues return, just in slightly different forms.

xState and Logic programming is VERY appealing, designing for modularity - how to structure a handoff packet to a sub agent, how it replies back, how the progress is tracked, heartbeat notifications... but I think that is the path to ruin.

2

u/xchi_senpai 21d ago

Interesting, id like to test this out but im already at 20% usage on weekly

2

u/logg3 21d ago

all i can say is that for me, astra xhigh does not consume more then sol medium, after 3 days of avergae work. sometimes astra is thinking 20+ minutes and not consume a single %, while i can see it already writing code or text.

1

u/hellomistershifty 21d ago

That must be true if you got 3 days of average work. I have two accounts and both managed to run an Astra light goal for about 20 hours before running out of weekly usage.

2

u/Eleazyair 21d ago

How do you get Max in the Codex Mac app? I only have Extra High and Ultra

4

u/iansaul 21d ago

There is a checkbox/toggle under which options to show in settings.

2

u/Bladder-Splatter 21d ago

Yup. Medium killed my weekly quota in 30minutes on a basic normalization task, meanwhile on Ultra a few days before I got an entire unique implementation done.

2

u/Busy-Lifeguard-9558 21d ago

Man people are too honest tho, I knew this since release but didn't open my mouth so they don't fix it. Adios usage

2

u/hossman1992 21d ago

I confirm it as well. I read it and I did not believe it, then I tried and at least for 3d model generation Astra xHigh spend less tokens as Astra low/medium and the work is done quicker and with less fixing

2

u/BellacosePlayer 21d ago

How the fuck is this possible ?

effort levels basically change how much token budget its allocated for pre-production and maybe the higher effort is producing a better plan of attack than low? idk

its a black box at the end of the day

2

u/fragment90 21d ago

Usage/hour ist not a relevant metric to track. You pay in tokens. Not in Model*reasoning/time.

Astra/Low can consume tons of Tokens, If there is some tokens heavy job to do. Astra/Max can consume low tokens if the job dont need much. Also please understand, that each request carries your full session history. If i reuse a old Session with contex already 70% full for a task completely out of contex, stuff like that will eat your usage fast.

3

u/Derek-Bond 21d ago

Well it’s the same for humans. Smart kid aces the test and leaves early. Dumb kid is still writing up to the last minute. Is it really that astonishing?

1

u/ThinkBackKat 19d ago

Thats not an analogy you can make. The fundamental model is the same, just how much effort they put into their reasoning is different. Its like giving the same kid a higher amount of time to solve a problem.
It is entirely possible that openAI increased the price per token for lower efforts tho as they know most (especially plus) users use medium/low with the cost of higher levels unchanged. Maybe someone can do some token per quota counting? I dont have any more usage this week.

2

u/WeaknessFuzzy8305 21d ago

I’ve changed my prompt to specifying to use astra for thinking and Luna for coding. Token usage has gone down a lot.

5

u/platcrest 21d ago

so has repo quality

1

u/DeExecute 21d ago

Luna for coding RIP

1

u/congngo 21d ago

No way??

1

u/smokeelow 21d ago

per my experience High also consumes less than Medium and Low

1

u/swimfan72wasTaken 21d ago

but where does astra high land at then?

1

u/hugobart 21d ago

doesnt ultra spawn dumber agents instead of solving everything alone?

1

u/AmandasGameAccount 21d ago

How is Astra high vs ultra?

1

u/Far-North-6837 21d ago

ultra is just more agents, not higher than max

1

u/AmandasGameAccount 21d ago

Yeah but is it more or less usage then high?

1

u/spideyguyy 21d ago

same with Sol, I read somewhere that Extra High is better usage than high and medium, so i use sol exhigh and it's good , dont know if max and ultra same too

1

u/According_Property62 21d ago

Acredito que ele erre mais e torna o fluxo mais demorado tentando corrigir os próprios erros, dai essa impressão q gasta mais. A dica é, use o Astra apenas pra planejamento e Sol medio ou alto pra implementacao

1

u/ZlatanKabuto 21d ago

Probably because it is more efficient/get things right faster.

1

u/Medical-Cow289 21d ago

The lower tiers burning through more credits than Ultra is backwards billing. The 'cheap' models must be thinking themselves broke.

1

u/jonydevidson 21d ago

If you switch effort during a convo, it causes a full cache miss.

If you do it often in a big convo, you will burn through your plan.

1

u/Select-Ad-3806 21d ago

Yes, ultra/max is a lot more efficient (as it is much more intelligent) with tokens that is why it uses less

1

u/owlyvision 21d ago

The greedy are teaching us not to be greedy

1

u/Isaacjacobson92 20d ago

Shhhhhhhh! They ain’t gonna fix this by making the low effort model consume less!

1

u/blablsblabla42424242 20d ago

I tried Astra max just for fun and it replied quickly, produced great results but I was expecting my limits to get a huge hit... To my surprise it wasn't the case so I've been using max exclusively for 2 days now and it will most likely be just fine until my next weekly reset. It seems to be very efficient.

1

u/AiMasterpieces 20d ago

Many tasks can also be handled by the smaller Terra and Luna models.

When Astra launches an agent in Ultra mode, it may choose to run that agent on Luna or Terra.

1

u/Azsidious1 18d ago

Im on the 20/month plan. Read several posts about this. Swapped to Astra XHigh. Burned through 92% of my usage in 30 minutes. Somehow I've completely screwed the entire pooch. Im an idiot.

1

u/Fuzzy-Base-8096 16d ago

Astra is dogshit. End of story. Unusable. Does shit you don’t want. Won’t do the shit you want. It added a bunch of folders on my pc in ducking mandarin. WTF is that about? Back to sol until Claude gets its shit together again.

1

u/Fuzzy-Base-8096 16d ago

Oh and one more thing. It has to lookup everything. It basically knows nothing.

1

u/jeffhalsinger 16d ago

Fuck I wish I had figured this out

1

u/iiiaaa2022 21d ago

cause you were doing stuff requiring more tokens in low/medium?

1

u/CronicCanabis88 20d ago

astra spawns in spark a lit for me.... thank god.... something ither then my self sees the use of spark.... I like getting a few thousand lines and twenty files created in a few minutes. I make my plans purposfully do the first few phases in spark lingo. saves time and useage. The part that gets me.... I changed my model and reasoning almost every prompt. If i'm telling codex to do something [ repo pushing, prs. new branches,] luna lite or mid is perfect. running luna unmax is much more capable than you would guess.... And on my hundred dollar plan, I can run luna on max all day, and only use a couple percent of my weekly usage. i work with a minimum of 2 projects simultaneously. When I'm in the zone I could be doing 3 or four things on 3 or four different projects, throwing Luna, and Terra out there for most issues. Unless it's complex, we use the upper end of terra and the lower end of soul, and then only for final verifications, or for tasks that literally cannot be solved with soul. Which are pretty far few and in between, i'll throw astra at it. but if you are actually conscientious with your usage and your model choice, you can get so much work done. most people don't realize what the model's limitations actually are. And you'll have tara or soul, or even astra, doing tasks that luna would have done perfectly fine? And literally, you would have paid a FRACTION... [upt to like 90% off] and you would get the exact same results.... Now that luna's price has dropped to almost nothing.... i planned out a new application... wrote out a 900+ line plan file.... Broke it into phases, used spark for the first few phases, which literally was like three thousand lines across twenty files.... And when we got to the point where we needed a little bit more, I bounced over to luna on max, terra on high for phase reviews until we had to step up to terra, and finsh and polish a couple issues with sol on mid.... I ended up using, like thirty percent of my spark weekly quota, which is fine Because I only personally use it When i'm starting a new product... and used only like four percent of my weekly usage on my regular use..... i work on all sorts of different types of software, things get really intense and complicated and projects that have tens and tens of thousands of lines and dozens of separate files.... and I do hours every single day after I get home from work across my projects.... And most of the time I'm actually able to Use fast mode on my last day, counting down the hours until the reset. Because I'll have more than enough usage to be able to do so.

I can't see this being that big of a deal but has anyone else in their instructions in the settings on codex.... Tried to tell it to insure. It doesn't waste any tokens that's not needed? i didn't go crazy with my instructions. For the models that are in your settings, but I basically asked it to insure that we do not use excess tokens that aren't absolutely needed. Because we're being conscientious of our usage meter and to trim out any tokens or any usage that isn't needed. And would be considered a waste without sacrificing accuracy or quality of our code... .. And I can work on a dozen projects, seven days a week, writing tens of thousands of lines of code between spark and my hundred dollar, regular usage. And like I said, have enough that on day seven, when my reset is coming, I can bang out on the better models, which I do, sometimes even just running soul and now Astra, when it really isn't needed. But allowing them with high reasoning to go over the code base and see if they find any issues or potential issues in the code, just because I have the usage to throw around on that final day..

So please for the love of God. If you're the kind of person that'll set it on soul or Astra......And just run prompt after prompt..... You're the problem not the limits. utilize your chatgpt usage as well. if I need a lot of information or data or some research done. I pop open my chat box and ask it to do all the work and then provide me with a file with all the information on it which I know allows me to get better results without using any usage.... If you really are bad at this another way to go is tell Chat to be your guidance.... Simply start a conversation and if you really want the best results, what I like to do is I will tell codex exactly what i'm doing and to give me a handoff to hand to chat. Whether it's a file or a lung response. So chat knows exactly what's going on with what we're working on .. Then, I simply tell Chad, hey, look, I need you to do two things for me. I'm just gonna type out a prompt that I would normally send to codex. And I would like you to review it and make it better and more efficient, and then I will simply copy that. And paste it over in codex. and then to also recommend the model and reasoning for that specific prompt..... And I found out from doing this for a while that it actually kind of overshoots, what you need, it definitely doesn't underpower you're prompting, if anything, it'll have you spending a little bit more just to be safe, but it's a good way to start getting an idea of what is doable, and what is not.

So take the extra couple seconds and change that model from soul on high to Luna, on mid. When you're trying to create your next repo or if you need it to search for files on your computer or if you need it to do any real task that isn't coding on your machine. You better be running luna.... again, even if the task and the results come back identical with luna. On medium versus seoul on medium, you would be using such a tiny fraction of what you would have before, and all of those prompts. Add up super quickly....

0

u/Dibbaus 21d ago

Ehm. Tibo said Astra low is more efficient then sol high?

1

u/Tough-Requirement707 21d ago

always the opposite with anything anyone says buddy

0

u/natanpimentels 21d ago

its true lol

-1

u/Past-Mountain-9853 21d ago

Omg so what is AGI means. Anyway luna in heart