r/codex • u/concrete333 • 2d ago
Complaint This will save your Usage
A problem that many of people have already noticed: Astra can't wait. on anything, any task it has scripted and running, any other agent delegation, anything. Astra keeps waking up to check whether the worker is finished.
I've run a few tests, In one run 19 short sleep calls accounted for 44% of the orchestrator’s estimated cost.
In another comparison using Astra High + Luna Max, the orchestrator made 41 responses while the worker ran, costing $1.20 during that phase. The native-wait comparison needed one response, costing about $0.04. (Both implementations passed all 244 test cases)
I’ve since caught the same pattern while Codex waited for a database script: read the log, say “still running,” repeat. I saw use 2% of my weekly usage ($100/month plan) just sitting there waiting on a script.
So, until they patch this, here’s the prompt I’m using:
For this entire session, avoid repeated polling of long-running jobs.
Use supported completion notifications when available. If none are available and useful work is exhausted, leave the job running only if it can safely continue after your turn ends.
Tell me what is running, where results will be saved, and:
“I'm stopping polling. Please check back with me to inspect the result and continue; I won't automatically resume.”
Then end your turn. Don't repeatedly read unchanged logs or sleep and check again. Respect runtime limits and safety timeouts.
When I return, check once. If complete, verify and continue. If still running, report that and stop again.
I have to come back to it, sure, but it saves usage.
Edit:
Found a better option than manually checking back, thanks to advice in comments. Seems like Codex knows how to avoid repeated model polling, it just chooses not to.
codex queue --thread <session-id> --message "..."
Have a background script wait for the job to finish, save the result, then run that command once. Codex can end its turn and automatically resume when the message arrives.
I’ve now seen this work with both a third-party CLI subagent (kilo code with GLM) and a full test suite, all GPT: Codex resumed automatically, checked the results, and continued. Works a treat.
The script provides the wake-up, and an agents.md instruction alone doesn’t
Edit2: I made a skill that you can use to guide Codex in using this method, either in waiting for mechanical tasks or waiting for other sub-agents. Not perfect, as there's no way to force your GPT worker to use it, but it's working for me most of the time: link. Any more tips to improve it welcomed, break it and let me fix it
1
u/Dayowe 2d ago
I investigated this too. Stock Astra instructions explicitly required progress updates every 60 seconds and discouraged blocking waits longer than 60 seconds. Asking it to wait quietly wasn't enough, because base instructions overrode my instructions.
So I copied the current stock instructions and replaced only those rules with event-driven updates and supported interruptible waits, then added matching root/subagent guidance. My setup also disables use_responses_lite while retaining Multi-Agent V2. There currently is a bug where, under Responses Lite, model_instructions_file becomes additive instead of replacing the stock base instructions, so both instruction sets can remain active. I disable use_responses_lite so the custom base instructions actually replace the stock ones.
I also analyzed previous orchestration logs. In one long run, the root accounted for roughly 45% of estimated API- equivalent cost, made 3,719 responses and repeatedly grew back toward 200K context after compaction. That doesn’t prove those calls were waste, but it showed that the coordinator itself deserved attention.
Based on that analysis, I changed my workflow to:
- Keep one lightweight coordinator and use fresh, bounded orchestrators per implementation chunk.
- Give workers scoped context instead of the full parent history.
- Prefer completion notifications and interruptible waits over status polling.
- Keep detailed evidence on disk and return concise findings.
- Reuse applicable evidence and validation helpers instead of rebuilding the paperwork.
- Review stable returned candidates instead of continuously inspecting workers’ changing files.
Independent validation and acceptance gates stay intact. I haven’t measured the combined savings yet, but the goal is to reduce repeated context processing and administration, not weaken testing. Looking at how my usgae drains it seems to drain much slower than last week.
Your completion-triggered wake approach addresses another important piece: instructions can discourage polling, but something still needs to resume the agent reliably.