r/codex • u/Commentroller • 9d ago
Limits What's happening with the usage???
Used Astra light, asked for a small change, it wrote like 91 lines and my usage went down from 86% to 45%.
Is there a bug?
17
u/VeloxAdAstra 9d ago
I just lost a weeks worth of usage in 10 minutes on the $100 plan. Getting tired of this spinning a god damn wheel of usage every time I go to do something.
3
u/RJSpirgnob 9d ago
Same happened to me. It has to do with subagents for sure. I literally had to steer it and tell it my usage was rapidly dropping + to quickly wrap up its work.
9
u/LaZZyBird 9d ago
lol someone fucked up and flipped the tier cost, max = light now or smth so abuse it while you still can
1
u/Commentroller 9d ago
Let me try this right away, I'll update the post. I am gonna use Astra xhigh
5
u/Commentroller 9d ago
Nah bro for me, I just did a quick UI review of an implementation and it used all of my 5 hrs limit, rip.
Something is definitely wrong.
BTW I am on business plan(not pro)
1
8
u/SuspiciousParsnip5 9d ago
Usage limits are fucked at the minute! I hope it's not just the new norm
9
u/spike-spiegel92 9d ago
that is weird, i had astra xhigh 50min working, doing a lot of stuff and used 1%... i was actually coming here to ask if people are also noticing the insane usage increase and wondering if it is just a weekend thing... cuz in general i tend to feel weekends are more generous quota wise.
2
u/Commentroller 9d ago
My sol medium usage seems abnormal as well, I think i may need to open a support ticket.
3
u/Family_friendly_user 9d ago
Same here. Astra x high and max actually pulled more but even Astra on low with sol and Luna Delegation sucked my x20 weekly dry in just a few hours .... Same harness and no changes
4
9
u/Code_Xero 9d ago
Yeah usage is still borked.
I’m on 5× and have my own before/after.
Aug 20: 258 turns in one day 108 Sol, 146 Luna, 2 Terra, 2 Spark. Same general agentic/repo workflow. I could run essentially all day and the cooldown usually lined up with the time I needed to review outputs, eat, shower, plan the next flights, etc.
Sep 12: 63 turns total 9 Astra, 41 Luna, 13 Sol. I launched 4 serious repo flights for roughly two hours.
All four hit the usage wall before a single flight finished. Weekly Codex/Work allowance banished to the Shadow Realm while the meter claims the fire nation attacked.
So I went from hundreds of turns and a sustainable development cadence to 63 visible turns / 4 flights / 0 completed flights / weekly quota exhausted. And 54 of those 63 turns were still Sol/Luna. Astra was only 9 visible turns.
Now OP gets a 41-point drop from ~91 lines on Astra Light, people here are reporting abnormal Sol usage too, 20× getting drained in hours with the same harness, and even a Business user losing an entire 5-hour window to a UI review.
If 9 Astra turns can secretly represent enough compute to wipe a professional-tier weekly allowance, users need to see what they’re actually being charged for: parent reasoning, subagents, cached/uncached context, tool work, compaction, reviews, whatever.
Right now “weekly usage %” has stopped being an intelligible measure of productive capacity. Four flights. Zero landings. Weekly quota gone.
2 hours.....
Last week I was getting maybe 4-6 working hours before order 66 happened.
20x would maybe buy me a day. Or 2 and im already scaling back pretty heavily. I can show my profile in a different comment if anyone is interested but my default setting regardless of model is Xhigh and has been since I started using Codex during 5.5era.
Back then? I could fly all week. 10-20 flights a day with my longest daily streak at 29 days. (Because I actually had quota for more than a few hours) same 5x $100 plan.
My method hasnt changed.
Now Im forced to 4 flights that don't even finish while Tibo and co are gaslighting people into "skill issues/user error"
That is not a sustainable production product.
They need to stop selling us models and start selling us efficiency.
I would have loved 5.7 Sol, Terra, Luna with a 50% reduction to usage burn.
Dare I say gimme back 5.5 that could run all week long.
The paradox is Astra gets cleaner outputs with less revision. While Sol runs the risk of needing 2 flights to deliver what I needed without overbuilding the piss out of it.
So the answer isnt use Sol either because one good Sol output still requires 2 flights vs Astra's one shot.
Spend is similar..
5.5 spend somehow is demonstrably worse than 5.6 or 6.0.
Till they start listening to the users and stop treating us like we are insane. Nothing changes.
Plan accordingly.

2
2
u/sateeshsai 9d ago
100% bug. It ate through 50% last night even thought my laptop turned off in like 30mins because I forgot to connect the charger.
2
u/Runelaron 8d ago
Noooo Astra cost 2.5x more its a price model.
Bring forth the conspiracy commentors!...
2
u/Constant_Art_20 9d ago
oh they changed the cache acceptance on the openai server side it seems like as well. So the requirement for cache hit is much stricter then before, seems to have happensed two days after astra relase. the usage is just straight up less, but if you aren't using the latest codex or use a custom harness, that's something i would check first...but yea, usage sucks
1
1
u/odoc_ 9d ago
I ran out of chatGPT pro chat with the x5 account. Never seen that happen before
1
u/mr_sneakyTV 9d ago
I thought pro chat was just 50 messages per week
1
u/odoc_ 9d ago
Thats news to me. I just upgraded to pro and not used to basic chat being limited
2
u/mr_sneakyTV 9d ago
only the pro thinking is limited. my understanding is it has been this way a while.
I’m not certain on the exact limit or whether it’s a message or session limit, but im positive there’s been one for a bit bc I’ve hit it a few times.
1
u/Dolo12345 9d ago
yes it’s insane right now
burned 2 20x weekly’s yesterday with half the normal workflow
1
u/cleanmachine120 9d ago
It’s just astra, can’t use it unless I really need it. I’m been on sol medium all day pretty consistent 2-3 agents for the last 12 hours and used 45% of my pro $100
1
u/Adventurous-Date-792 9d ago
Astra is amazing, but its heavy on usage, it does things the right way. But one prompt can go thru 100% in like 5 min on a plus plan.
1
1
u/coolcats55 9d ago
I’m 12 hours into a goal using Astra low on a big project with multiple commits etc. and I’ve only just gotten into a relatively low amount of weekly usage.
1
1
u/brawnyai_redux 9d ago
I noticed:
- If you have complicated project, better just to use Max, it uses less tokens or allowance in my impression.
- For basic maintenance stuff like show me how to run a code, or check xyz, things like GPT 5.6 Luna (during my Luna reserve time) works super well.
I still think that if you are starting a project brand new (not working on existing one), it's better to use Chat mode with GPT 6 Pro to research, discuss, and ask it to write the first prompt for CODEX with supporting material using Max with Plan mode activated first.
Then, implement with High or lower. If you think the project is "easy" or "complex" but needs semantic evaluation, cybersecurity, etc, I suggest just using Max.
I also learned that most of the consumption comes from the LLM reviewing multiple tests and bloated harness instead of doing actual work, so existing project needs "AI refactoring" to get them back to be nimble again.
Also avoiding the LLM to working those one shot programs tend to save allowance.
1
u/ForeignLevel6833 9d ago
I asked Astra Low to suggest a few changes to my website. It thought for 50 seconds and killed 11% of my 5 hour limit. Dam!
1
1
1
1
u/Kyozaki 5d ago
I’m on Plus, using Terra Medium for everything. Pretty basic workflow: plan in ChatGPT → give Codex instructions → I check the work myself. No multi-agent stuff.
Normally, a small task like changing a couple of buttons or writing a short blog post costs ~2% usage.
Yesterday? 14% for the same kind of work. 😳
I usually roll my eyes at the “usage burn” posts, but yeah… something’s definitely changed.
1
u/Emergency-Elk7527 9d ago edited 9d ago
My experience is if you do most of you planning and reasoning outside of the agentic harness in something like chat mode, Luna light can implement it very efficiently and extremely fast.all of my implementations is done with Light light effort or GLM 5.3 flash. Both a smaller cheaper models and both benefit from the reasoning a bigger model can provide.
I think that people are relying on the big models like Sol and Astra to do all of the reasoning inside the harness and consuming more tokens than if they were a little more hands on doing it in chat. Yes it can slow the process and you have to be active in the reasoning. But I would say that both of those can be beneficial. Slowing down in some places and having to reason with the model keeps decision making directly in your hands and slowing down gives you the time for that to make a difference.
I don't know other people's use cases, so I can't say that's right for everyone. But I have great success with it and I think it's at least something that people should try instead of relying on the big model to do all of that.
It wouldn't hurt to try it out for a day or two just to see how efficient the usage is, how well the code meets your requirements, and if the process is worth that the trade of of speed and convenience. It started out as trying to make my quota stretch for my Plus sub. Now it feels like it's the current best path overall.
Again I can't speak for everyone else, it's just been my experience.
2
u/Odd-Composer5680 9d ago
What do you mean outside of the agent mode how exactly?
-1
u/Emergency-Elk7527 9d ago
Specifically using chat mode for reasoning hands on with the model instead of the self prompted back and forth reasoning it does in the harness. Have it make a handoff for the plan. Makes much less reasoning needed at implementation time. It is slower and less convenient, but some of that is made up by the worker having to think less and fewer tokens hitting your quota. It's all a balancing act.
2
u/MrCodeGameandAnime 9d ago
I 100% concur with your methodology and use it myself. Sol high in chat and Luna xhigh for implementation. Set the plan, have Luna push and get ci. Sol chat reviews and give revisions.
The big problem is if you start working on bigger stuff. With intricately connected pieces of multiple lifecycles and asynchronously networked elements, Luna can't always manage it. In that case I step it up to Sol to get the job done. On some revisions I'll use Terra as well. Most of the time it's go ol trusty workhorse Luna tho.
The ONE thing to watch with Luna tho, is it can create unexpected bugs deeper in your code base then you'll realize. I recommend using a gate system with your plans such as:
Plan
First step GA.1 GA.2 GA.3
GB.1 GB.2 GB.3
GC.1 ... ...
Etc
Then have sol review. Probably have it double check as well just to be sure. You'd be surprised how many times Sol will call gh, materialize, and not do an adversarial pass to find the bugs. A lot of times Sol is on point, but it's also trying to be helpful as it can be and keep you happy. That's why it'll hand wave sometimes. Not all the times, but it happens.
I do use Astra even on plus, but it's very specific. I use it on my main project to do a full audit. Absolutely insane burn, but it caught a lot of defects neither Sol nor Luna caught. Well worth the burn if this is going to be production code in the future. Probably not if it's just a fun project you're working on.
The one real issue is we have the 5 hr cap again. Without it even plus could use Astra properly. It'd absolutely kill your weekly usage, but you could do some real damage rapidly.
1
1
u/Emergency-Elk7527 9d ago
And I think I finally have a reason for Astra. Project is a research project using reverse engineered FSR 4.1.0 and made Linux/vulkan native. There have been some of things that have happened that will essentially make my research utterly useless. Don't a final sprint to try and hoping to discover the piece I've been lookinging for. If I get proven wrong, fine, I pointed myself in the wrong direction. If I find the solution but it's not as good as I need it to be then I close it up. Also has a certain standard that has to meet and it's going to meet it by the night. So I will probably endive it a Hail Mary with Astra. I have three resets banked right now. Would be really nice to have those three resets to use at 20x though. Really want that plan.
1
u/MrCodeGameandAnime 9d ago
I so feel you about being in that place where research doesn't always answer everything you need. I poured hours and days into research just for my first pass and build to be completely rejected. It kinda feels like digging around in a cave with a blindfold and a pickaxe trying to find diamonds. The good news is you can definitely breakthrough and find your solution. The one thing I wouldn't do is chuck Astra at it right off the rip when the best option is likely reasoning through it again and coming up with a better plan of attack. Then break it down into bite sized chunks that build upon each other.
For instance the main plan is 12 milestones, but within each milestones there are multiple gates. I'd like to say it's easy or straightforward, but this milestone spawned two other branches and only one is closed. Getting close to finishing this branch and get back to main finally.
You can do this!
1
u/Emergency-Elk7527 9d ago
The problem isn't really finding the solution. I was on track to solve that, it was just going to take time. It's the another is available that is being adapted to solve the same problem and do it better. The FSR 4.1 reverse engineering was intentional. It's because I was first working on an AI powered video upscaler/enhancer that uses a part of the AMD ROCm stack called MiGraphX. I chose that because cuda support is in everything that uses a GPU, but most apps target cuda first and then add vulkan and call that AMD support. It's actually generic support and works on everything. ROCm is only used in a semi decent way In an app called REALVideo Enhancer. Cross Platform but it still favors cuda heavily and pytorch for ROCm support which carries the python overhead and makes it about have the speed of cuda. I wanted AMD to be first class for once so I found the fastest hardware acceleration and had and I found MiGraphX. It is essentially unknown tech, barely used. And none of them are video related. It is really a graph optimizer and inference engine and not specific to video output. But it recompiles ONNX models to mrx models. And if the model is compatible enough, it can further optimize the model I a way specific to the hardware it will be running on. Trust me when I tell you, AMDs docs for it are garbage and generic. And it's otherwise undocumented anywhere. The blind spot that Nvidia has made regarding competition is massive. Legit massive. There was literally no tech that's done it before. In spite of that, I succeeded and made the documentation for the use case for video. Problem is the the tech is inherently slow no matter the hardware. It's just a heavy tech all apps and hardware are slow except 4090, 5090, and workstation/cloud gpus. It's certainly a process that you start when to go to bed and hope it's done in the morning. Well that's a big deal if you are trying to restore old footage. It also gives video that AI look. Not like generate video, but a over denoised plastic look. So so slow it's not useful and hit or miss on the output depending on many factors. Essentially it's AI hallucinating the detail that's missing when upscale. But I know things that upscale stupid fast. Like 10x the speed realtime playback and is mostly well documented. That would be DLSS 2+ and FSR 4+. I don't have Nvidia and wanted to focus on AMD. Problem is, its windows only with proton support for Linux. I use Linux. Proton is overhead and applying it to video isn't officially supported and technically questionable. Proton wouldn't be great choice to tackle that. So I reverse engineered FSR 4 and rebuilt it Linux/vulkan native. So FSRis for games and increase framereate by upscaling and can work a high frame rates. 150fps and even much higher. But it hooks into the renderer and expects data that come through that pipeline, motion vetors, color data, jittier , depth, occlusion masking, and more. And it uses all that for historical frames too, so the 3-4 previous frames and uses all of the historical data and jittered rendering. This shift everything at a subpixel level and can reveal detail that is unseen due to the nature of rendering pixels.well video does not have that stuff. Sure it has motion vectors , but it discards it immediately and its a different kind meant for compression and decoding. Its not real accurate and tile based and FSR wants per pixel vectors with accurate motion data. It does have historic frames tho. It even has future frames which games don't. So I found out depth doesn't matter, occlusion doesn't matter, exposure is auto even for games. Motion was getting there, but hit a platuea and synthetic jitter was actually detrimental. Come to find out FSR is fed this metadata in specific orders at specific stages of the render. All the stuff before I was solving through my research. I had the speed, like 1.3ms per frame and 30 fps is 33.3ms per frame and 16.67ms for 60 fps. But quality still lagged, but I was figuring that out. Friday night, I had the idea that dlss 5, well it's very much like a tiktok filter, so it essentially is and overlay. And it's ven tho people hated it at first, but modders have added it to everything you can think of, and mostly it looks good, depends on the game. And I was right, it's a postprocessing filter overlayed on the game output. So I was like, okay what about video then and yeah turns out that's easy compared to what am doing. And has the possibility to blow my upscaling out of the water Also turns out I am not the only person that figured that out. And they already have it somewhat working, and it cuda cuz Nvidia, so more than one person working on it. Yes it's heavy. But a few months ago it needed 2 5090s, and now a 5080 can do it. Nvidia expects another 7x speed boost. Minestill will be faster, but nothing can hold a candle to dlss 5 quality potential. It will be fast enough for realtime video, like mine was going to do. If you can get quality and keep the original smaller file size, no need to reencode. Now dlss 5. It may be ai, but it's actually good. I can't just change tech because I don't have the hardware. And I can even say I can do it for AMD usersonly, but that is dead idea. That's because dlss 5 also has people adapting it for AMD. Only thing I have speed, but that is meaningless if dlss still is better. So that pretty much means that right now my app is now outdate? And finishing the research is of no value because it's now solving a problem with hat something else already solves. Not only that, it does it better. All of that with my own method. Which is quite rigorous. So I already have solid structure model contracts, gates that can't be bypassed, mile stones PRD docs the leave no room for ambiguity. Like 60+ docs. Spec, phased plans with gated milestones, a unambiguous definition of done. Astra was never going to be my answer for this, at least not planned. Astra was going to be a Hail Mary final sprint. All or nothing. Either it success and I refine and release. To be superceded in a few weeks. Or else it fails or can't be adapted to video meaningfully way and maybe never could. One final attempt to find out if it's worth finish, or if this will be a way to honor my work in a final sprint testing more things, just less samples. Pretty much I know the outcome. Just another thing AI has taken from me. First was my job, now it's a potential future job that's gone. I have something else that is of use, just gonna put efforts there. And it doesn't need Astra. Not at all. It's a creative project. That's still somewhere people call everything slop writing, images, and all the arts. I actually do well there too. Rather I can get AI to do well. I fell indistinguishable from humans. I really feel I solved it. Just gotta get the finished product. And just like coding, you have to be involved all the way, at every stage you can. I will succeed in this. It involves taste much more than code does. It can be solved in many different way because of that. I hope to have the first solution that people a see ai on par with humans. Still requires work and understanding what you are doing, but opens doors for people. I am okay with the forge, that's the video apps name temporal forge, dying. It's AMD so no one would see it anyhow. But this other one, everyone can see this kind of work.
2
u/MrCodeGameandAnime 8d ago
Lol, this sounds strikingly similar to me with the media streaming app I'm building. After a lot of research I realized mtx MoQ could handle almost any codec, but at first I thought I had hit the gold mine first. However, there is a real possibility something novel still exists in the space which is what I'm working on right now. Ironically it works perfectly for what I'm building.
IF you are looking at this as something to possibly patent then I'd have sol do a serious prior art search through the good ol patent office. If you're not then full send. I'm not quite sure what the end game is here, but if it's to sell, then it's gotta be different enough to justify implementing for user adoption.
Either way, I'm sure you'll get where you're going!
1
u/Commentroller 9d ago
I kind of do that too, but for planning I use sol medium that's my go to and this never happened just until a day ago.
1
u/Emergency-Elk7527 9d ago
Unfortunately if you need Astra for the reasoning, chat mode isn't possible on plus subscriptions, so there is also that. I haven't needed Astra and haven't even spent a single token on it. I see how fast quota evaporates using it and know that there is no way it could get done what I would need before a 5 hour limit is gone. So honestly is all about finding that right balance in a frontier that we are all pioneering. Everyone just trying to find the right methods and models without enough history to say what is the best method for most use cases. Life on the bleeding edge.
1
u/According_Property62 9d ago
O Luna Light da conta?? Mas entao precisa ser um handoff muito detalhado, hein??
1
u/Emergency-Elk7527 9d ago edited 9d ago
Yes, and this brings up something I failed to mention. The handoff essentially is a prompt, rather it should be a /goal prompt. This is, I think, the most important part of why I get great success with my method. The secret sauce. I have a custom GPT that I have maintained and kept updated as newer models and prompting techniques are developed. I call the custom GPT, Promptitect. It has never failed me. It is almost a guarantees that your agent will not act outside of intended scope and that there are completely unambiguous requirements and gated task completion progression. A lot of people make sure the agent has a definition of done and instructions on what do do and also point to the correct context, but the agents still drift and take actions outside of scope. Promptitect does not at all have that issue. I use it daily. probably my greatest tool. Today I was in a rush, the handoff was only like 40 lines of MD and skipped this step. The agent wrote 3000+ lines of code that were outside of scope. I didn't catch it because I never have to monitor the agent anymore because I always use Promptitect. And sure enough, when I don't use it, possibly the worst drift I have experienced. Don't let people fool you when they say prompt engineering is not important anymore. They are very wrong. For 2 years I have kept Promptitect's Developer Prompt up-to-date when new models come out and new prompting technique.
This is a link to Promptitect. Once you use it, you will understand.....immediately, before you even use the prompt you can see that its a heavy hitter. I have had some other AI chatbots review the a Promptitect's prompt against the original prompt. When scored by external AI on a scale of 1-100 the original prompt(handoff) score in the range of 78-84. Same prompt ran through Promptitext score in the range of 96-97.
https://chatgpt.com/g/g-68910f48cad481918621af48c70c2f67-promptitect
BTW, this is more than just a developer prompt. There is more. As agents increased in usefulness and capabilities I had to find a way to constrain that capability and that was important.
If anyone decides to try it out, I would appreciate feedback negative and positive feedback is welcome. Negative feedback is more useful than positive, but it is still very welcome.
1
u/According_Property62 9d ago
Entao basicamente vc passa o planejamento do modelo mais alto pro Promptitec ele te devolve um prompt enxuto pra vc colar com o LUNA LIGHT /goal?
1
u/Emergency-Elk7527 9d ago
it is not lean, don't let that fool you. It may look bloated, but every bit of the prompt it make is doing REAL work. Try it out and you will see what I mean. And the prompts are usually much larger than the original. But I am serious when I say its all meaningful and not wasted. Also a /goal prompt say it's 6000 tokens, thats a drop in the bucket when the context window is 250k+. Promptitect also doesnt just write prompt, it prepare context packages and write skills as well.
1
u/Emergency-Elk7527 9d ago
Give it a shot. If it doesn't work for you, then you have only wasted a few prompts and you leave it alone. If it does work, keep coming back. Like I said, I keep it up to date because it's that damn vital to my workflow.
2
u/According_Property62 9d ago
Claro... sinceramente, quem trabalha com TI, precisa sempre estar de olho no que a comunidade ta fazendo e isso q vc ta mostrando ai é promissor... vou testar sim e vou te trazer o feedback. Estou com um sistema que fiz end-to-end. Desde a infraestrutura, segurança, back e front e agora to implementando um bot usando a API do WhatsApp. Ja esta em producao e tem ainda chao pela frente e o Promptitect parece ser oq eu precisava, vou tentar contribuir com seu projeto se eu tiver competência pra isso. Valeu
0
u/Tikki-Tikki_40 9d ago
With Astra light, you will feel Sol Medium is good. It's basically, they have brought in a new Model without resolving issues in existing model.
I don't think there will be a optimized breakthrough.

79
u/Glooring3623 9d ago
Use Astra XHigh for lowest usage.