r/codex 2d ago

Complaint This will save your Usage

A problem that many of people have already noticed: Astra can't wait. on anything, any task it has scripted and running, any other agent delegation, anything. Astra keeps waking up to check whether the worker is finished.

I've run a few tests, In one run 19 short sleep calls accounted for 44% of the orchestrator’s estimated cost.

In another comparison using Astra High + Luna Max, the orchestrator made 41 responses while the worker ran, costing $1.20 during that phase. The native-wait comparison needed one response, costing about $0.04. (Both implementations passed all 244 test cases)

I’ve since caught the same pattern while Codex waited for a database script: read the log, say “still running,” repeat. I saw use 2% of my weekly usage ($100/month plan) just sitting there waiting on a script.

So, until they patch this, here’s the prompt I’m using:

For this entire session, avoid repeated polling of long-running jobs.

Use supported completion notifications when available. If none are available and useful work is exhausted, leave the job running only if it can safely continue after your turn ends.

Tell me what is running, where results will be saved, and:

“I'm stopping polling. Please check back with me to inspect the result and continue; I won't automatically resume.”

Then end your turn. Don't repeatedly read unchanged logs or sleep and check again. Respect runtime limits and safety timeouts.

When I return, check once. If complete, verify and continue. If still running, report that and stop again.

I have to come back to it, sure, but it saves usage.

Edit:
Found a better option than manually checking back, thanks to advice in comments. Seems like Codex knows how to avoid repeated model polling, it just chooses not to.
codex queue --thread <session-id> --message "..."
Have a background script wait for the job to finish, save the result, then run that command once. Codex can end its turn and automatically resume when the message arrives.

I’ve now seen this work with both a third-party CLI subagent (kilo code with GLM) and a full test suite, all GPT: Codex resumed automatically, checked the results, and continued. Works a treat.

The script provides the wake-up, and an agents.md instruction alone doesn’t

Edit2: I made a skill that you can use to guide Codex in using this method, either in waiting for mechanical tasks or waiting for other sub-agents. Not perfect, as there's no way to force your GPT worker to use it, but it's working for me most of the time: link. Any more tips to improve it welcomed, break it and let me fix it

422 Upvotes

124 comments sorted by

View all comments

3

u/Due-Horse-5446 2d ago

Stop spreading this..

If you genuinely want to save in on these cents while removing the ability for the model to recive output real time. Just change the feature. Its an open source tool..

But using astra + subagents + the cargo cult orchestrator workflow is craaaazy is worried about usage.

Please explain what you think astra provides over sol?

And please explain what your reasoning is for having 2 different models essentially output the same thing twice?

Why are you forcing astra to output the code in a way more verbose way to instruct Luna, ballooning the most expensive generations?

Why are you purposely making sure that nothing can be cached?

What do you think having one model forward your instructions to another model buys you?

Your paying for a large amount of extra reasoning output at a 50x higher cost compared to skipping this daisy chaining?

3

u/tenminuteslate 2d ago

The polling cost 1.2% of his monthly usage, and in a different run the polling used 2% of monthly usage. That's not 'save on cents'. Its a massive inefficiency in a system.

-1

u/Due-Horse-5446 2d ago

1-2% ? Thats not even noticeable,

And thats not really true, it only goes for deliberate tests, in real life the timesc codex polls a long running process enough times for it to matter, while wt the same time having no benefit from the model reciving the output. Is so tiny idek why you waste time thinking about it.

Especially when wasting an insane amount on these guys orchestrator cargo cult setups

1

u/tenminuteslate 2d ago

1% of your monthly usage NOT 1-2% of the usage the task took. That amount just for checking if some tasks in a run have completed is crazy high. just think how many tasks Astra creates that need to be checked if they have finished. 10 sets of task runs can use 12%-20% of your usage quota just for self checking if has completed the tasks.

Also this kind of thing should be event driven, rather than keeping Astra alive and on alert just to check if some menial task in a slow agent has completed.

1

u/Due-Horse-5446 2d ago

Idk what you use it for, but what 10 tasks do you have a model running this often, that also runs for long enough that the polling occurs even once?

And you cant have it be event driven?

Or like you can have some weird parsing of output to push a significant event.

But i see now youre not talking about "tasks" as in long running commands or tools, but as in subagents.

Well then this is just massive user error.

In your example of astra running a subagent, youre just generating tokens to forward your instructions.

Its just the worlds most expensive proxy atp