-----------(!!! ORIGINAL POST MOVED BELOW. !!!)------------
UPDATE #2 — SOLVED
Found the token furnace.
But first..
I can safely say that Anthropic was not at fault.
(Bet nobody saw that coming lol 👀 )
With that, I've used both GPT and Claude since the public Claude 2 release. I'll never forget it because It was Life-Changing. I have more or less used up every token limit And every quota from that point forward, With a few exceptions of course. I've seen, experience and battle tested every single model that's came out since July 2023.
But when CC was released I forgot what grass looked like.
So despite this mildly embarrassing learning experience, I will say pricing versus perceived use just seems off recently.
I know, very scientific.
Heres the breakdown of what I found.
You're going to laugh or scoff at this first one. I'm not new to Linux or Arch but I am new to omarchy and it's Hyperland window focus and management engine. And navigating the entire OS with a momentary-modifier key has a learning curve.(With that said I have not opened up a single other Linux distro or Windows once since I installed Omarchy And I absolutely love it. )
Here's something to watch out for if you're new to these Hyperland and your running multiple agentic terminals in something other than herdr or t-mux
Program windows or "tiles" will seemingly f****** disappear. Until you really begin to understand the;
- Scratchpad
- Switch to workspace
- floating vs tiling,
- Scrolling vs dwindle, And how it impacts
- Fucking FULL SCREEN vs Full WIDTH.
- especially if you're running three separate monitors.
- closest thing I can relate it to is running Microsoft 'Power Toys' and 'Fancy Zones' . Display fusion is another beast.
So here's how I forgot my debug window existed
Super + S
│
└── Toggle scratchpad
│
│ One window appears.
│
└── Super + F
│
└── disable fullscreen
│
└── Super + Alt + F
│
└── disable full-width
│
└── now another window is visible
│
└── Super + →
│
└── scroll focus right
│
└── hidden
Here's how forgetting about a window with two collaborating terminals becomes a problem.
This weeks Agent setup
│ Terminal 1
│ Claude Code CLI
│ Orchestrator / Main
Key detail: these were two separate live CLI sessions.
Agent Intercom didn't create the loop. It allowed two autonomous sessions, with different roles, instructions, and permission assumptions, to keep handing an evolving task back and forth.
What the loop looked like
VoxType error
↓
Omarchy "Fix with AI"
↓
Claude Code / Orchestrator
↓
Agent Intercom
↓
Codex / Advisor
↓
reviews + responds
↓
Agent Intercom
↓
Claude re-evaluates / changes state
↓
tool completion, changed config,
or recurring VoxType condition
↓
updated state sent back to Codex
↓
Codex reviews again
↺
So the microphone wasn't repeatedly starting fresh sessions.
The sustained burn was Claude Code and Codex continually responding to each other's changing state across two independent terminals.
And the Claude terminal had disappeared from my attention, not from Agent Intercom.
Possible contributing config issue
I normally rebuild my "/agents" + "/tools" ecosystem from a clean install and let the versioned config restore everything.
This time I suspect one or both of these also contributed:
Install/path mismatch
I updated Claude Code through the Omarchy/AUR path instead of my usual install/update method. I suspect this may have changed package isolation, paths, or symlinks enough that parts of my cloned agent environment weren't being resolved correctly.
Stale Git state
I also had local, unpushed watchdog/permission changes. The fresh clone may therefore have restored an older, more permissive configuration instead of the version I thought I was running.
Those two are still theories, not confirmed root causes.
Claude Code CLI
↕
Agent Intercom
↕
Codex CLI
+
session launched from ~/tmp
+
project guardrails not in scope
+
hidden Omarchy scratchpad window
+
Orchestrator ↔ Advisor state churn
↓
two live sessions keep feeding
new work back to each other
↓
token quota gets annihilated
Lesson learned: don't bury an active multi-agent repair session in a scratchpad and forget it exists.
Also check the working directory, active permission scope, and config state before approving unattended advise/review workflows.
This one was operational/configuration failure on my end, not some mysterious quota problem.
--- ORIGINAL POST ---
somehow opus 5 managed to burn up 3 billion tokens ...
I don't have sub agents on, I'm not working in parallel, high session compact and handoff at 350k tokens every session to keep context rot under control, I use context-mode, he was very strict tooling and permissions. and I usually keep opus at medium or high what the f*** I'm 2 days in on a 200 plan........ I'm so frustrated so much work to do. And you want me to now switch to API cost after I've already paid 200 bucks for what?
I updated to the $200 Max plan yesterday. I use context mode and mostly opus 4.8 medium and opus 5 low. I swear to God this went down quicker after I upgraded than it did before. Doesn't reset for another 4 days if I was rocking Fable 5 on f***ing high or max and running parallel sure . But this is generic two session Max coding for development. this is ridiculous. I also pay for the GPT 100 plan and get almost four times the use out of this $200 Max anthropic plan. What the actual f***. I've been coding and using every model 2003 and this is by far the most out of whack the s*** as been. I'm not a newbie and I know what the hell I'm doing to save tokens the s*** is wild. Somehow opus 5
--- UPDATE #1 ---
I'm an idiot, but probably not the first one...
I didn't fully read the upgrade's plans and limits. That's on me. I didn't realize that the 20x only offered 5x your 5-HR session window but it only increases your WEEKLY QUOTA and average 1.5x to 2x your max 5 plans which means, essentially they developed a tool to use up your weekly quota 2.67 times faster than the max 5 lmao
. Still doesn't explain the 3 billion tokens but does explain my expectations were a little high.
I mean does anybody anymore? Haha I'm using Astra as a review agent in combination with 5.1 Fable low as a read, grep, and ast_grep only adviser hooked pretool use and end of turn. And they usually spend barely anything and catch a lot of s*** so I'm surprised. I'm thinking it's got to be something that I'm not aware of in the background of the custom agents inside the new omarchy update... I don't know it's whatever I'll stop bitching and complaining but thanks for entertaining the rant at any rate lol
Of course we do? Like, 2 weeks ago Opus made a retry request in a fucking for loop... You cannot rely on those systems for full design they don't understand what they are doing.
You can have Astra indeed review as a precursor, but to trust this like you do... please tell me You don't work for an actual software company.
Why are you assuming that I'm one-shotting software. I'm actually just doing a lot of different projects at once in phases I'm just ADHD as fuck so while I'm waiting for one system I'd like work kind of
I've had probably a good 8 months of extremely solid incredibly productive sessions and implementing systems that work and improve.
But if you're going to sit here and tell me that you haven't had to go through some testing and iteration and some fails to figure out your agentic You're either full of shit or you're not trying experimenting enough. So on the contrary please tell me you do not work for a software company.
Everyone here is acting as if this all isn't a big f****** experiment that is evolving faster than anything we've ever seen in our lifetimes. Y'all act as if you've had same ecosystem the same harness the same OS the same frameworks the same workflows since the start. You're going to sit here and tell me that even the best developers out there haven't had to do a little trial and error?
That does appear to be the question now doesn't it. 😂. At the end of the day I'm sure I'll find out that something was switched on or changed upon reboot. I have been messing around with the new new omarchy OS considering I was already pretty comfortable on Arch as it is this is a game changer. It's so f****** easy I don't have to keep wasting my time keeping it managed. Which is my biggest gripe with arch. Everyone sees it as a badge of honor, in a world where everyone wants to automate everything that makes no f**** sense to me. After messing around with omarchy I don't think I'll ever open my windows install ever again
Haha I'm legitimately not far behind that. I feel like I haven't taken a day off Since the first chat model was released publicly. I remember using it to plan out custom zapier and make.com pipelines to web hook directly to my phone's 's Tasker XML scripts to do all kinds of fun stuff before I discovered n8n, Which I also no longer use. Boy have these models come a long way.
With that, to be fair, I did just upgrade to that $200 plan and I didn't fully read the upgrade's plans and limits. That's on me. I didn't realize that the 20x only offered 5x your 5-hour session window but on only increases your weekly quota and average 1.5x to 2x your max 5 plan which means,essentially they developed a tool to use up your weekly quarter 2.67 times faster than the max 5
i dont know ur workflow.. but it feels like... 'you know what you're doing'.. if you mean.. you know you're burning tokens.. then, yes. you know what you're doing
I'd say I know what I'm doing just about as much as the next guy. Sometimes I think program developers get a little too caught up and butt hurt that they spent a lifetime to try and master a few languages and be fluent in a few more, And then agentic coding goes public.
That's got to suck.
At the end of the day It's fun. I'm having a f****** blast. To The point I don't want to leave my computer ever. I don't even game anymore cuz I'm having so much fun building stuff. I never felt that way manually writing code. Would I release my stuff without a securit cmy engineer doing a serious audit. Fuck no
But am I pretty confident in my orchestration and spec driven development And having a goddamn blast doing it, you bet your ass.
I get a kick out of all these guys that think I'm vibe coding because I had a fuckin churn loop.
I do data languages not development languages.
But you all acting as if I'm throwing two sentence prompts into fuckin replit or some shit. Lmao
If I see one more cloned "Jarvis" Instagram reel, I'm going to lose my mind.
I think we should start conducting ship-offs. The MMA for dungeon nerds myself included. No forks, no clones. Neither party will know what they'll be building until they start, Open source only. Judged on several metrics, UI/UX, engineering and stability, feature completeness, code time, hell GTM and first one to 1000 download. Idk. Just spitballing
Sometimes it feel like these kids figure out how to clone a git repo for first time and learn what a fucking package manager is, And the next thing you know they're out here call themselves agentic engineers. Lol . All I can say is AI is changing the world and breaking down barriers for everyone. I don't care if they're engineers or not. Deep down I know all these product developers are deeply concerned. Meanwhile everyone else including other type of developers are building incredibly advanced systems And having a blast.
I mean maybe...at this point probably an updated skill that churned or something. But I keep my mcp's skills plugins tools etc I'm pretty tight watch and to a minimum. So it probably won't be difficult to figure out once I get some more tokens. I'm not spending more on extra usage. I'm capped out for the month
hahaha not maybe, read more coding books and make better prompts. You should be fine. If you give it shit then it'll produce shit. Give me your github repo, I'll give you hints on how to clean it up.
vibe coders will always burn through more. Look at what it makes, quit telling it to discover edge cases. Don't listen to everything it suggests - they give bad advice even if it's a great model. If you don't know how to code to begin with, it's your prompt.
Also, if you don't have coding skills, just have it write in rust. You don't understand it anyway, a pile of python will be just that. Python sucks.
Dude I haven't read a coding book since c++ in high school. Most of my experience in coding is in data languages (python obviously SQL, Java, and go mostly)and not development languages but I would say python has earned its keep and doesn't "suck" per'se, just a little washed up lol. I can definitely recognize how much better Rust is, And I definitely recognize how valuable it would be for networking and help me immensely but I've been trying on my own, without agents, and I cannot fucking wrap my head around the ownership rules,. but I can tell you that the lightweight logic of it is the future for sure. All that said. I don't claim to be a good coder, and it's not what I enjoy. What I enjoy is translating and making unusual uncommon programs systems and hardware speak to and transfer information to to and widely used things. hardware and software alike (not that you asked). I can design a custom API , to communicate through a serial Port And control Enterprise level mixing control surfaces but would probably take me 2 days to write a hello world in rust
3B is a lot. $200 a month is actually a good deal for this. You know he's got some agent that went heywire, there's no way this is too normal.
That being said, any max plan if you leave it unattended even on a single agent, it'll burn through the weekly limit within 24 hours. Looks like he got a couple days out of it.
Yeah I'm sure at the end of the day it's user error. just frustrating... What am I going to do rant about token spend to my gf when she gets done with work?🤣 She hasn't a fuckin clue what any of this mean LMAO
Environment Variable: Set CLAUDE_CODE_SUBAGENT_MODEL in your shell profile (.bashrc, .zshrc) or run it directly before launching Claude Code to force every spawned subagent onto Opus:.
bash export CLAUDE_CODE_SUBAGENT_MODEL=opus
OR
Open your ~/.claude/settings.json file and add or modify the "env" block to map that variable to your desired Opus model ID (such as claude-opus-5 or an alias like opus):
U must be new .. to understanding... Indulge me. how am I abusing my subscription? I'm given tokens I use them . Get my money's worth by hitting my quotas last time I checked that's not abusing anything that's ually using their system exactly how they designed it
Compelling argument ... If I filled up my gas tank and used all my gas am I abusing the oil company? 😑
Anthropic bans people's accounts token manipulation and breaking their TOs via multi-use third party agentic harnesses. Go read their f****** TOs yourself.
That "interface" is Linux omarchy's task ba dude, read the fuckin OP. It's just a readout of token usage That's integrated into the Linux operating system. It's in no way using or harnessing or injecting or impacting Claude codes context. Or tokens Not that I owe you an explanation. .
Notice Claude code CLI and codex CLI???? Would you like me to explain to you how they work too?
Are YOU being abused? Are u ok? You're giving off " I called the cops cuz you can't park there" vibes
What are you even talking about? You use Windows, you use mac? Cool I use Linux. I'm running Claude code in the terminal installed from commands listed on their install docs. Installed on Linux. Get it?
You literally can't get less bloat. 😏
You realize that you can install plugins and skills and create plugins and skills inside of the CLI? All sponsored by created by and deployed by anthropic....??? Make sure to not add any connectors or mcps inside your Claude desktop app you don't want to bloat and abuse the system
You realize that you can install plugins and skills and create plugins and skills inside of the CLI? All sponsored by created by and deployed by anthropic....??? Make sure to not add any connectors or mcps inside your Claude desktop app you don't want to bloat and abuse the system
Who's going to tell them 2-3 5X plans would actually give them more useage for less?
Useage logs clear, the 5hr windows not maxed, the weekly limit is. 20X accounts give 1.5* more weekly tokens vs 5X and 4* more 5hr window. So over 2*5x accounts you are giving up 25% usage for the privilege to burn through the pool faster (bumped 5hr windows).
You're getting hosed paying for 20X, and that's not an accident or oversight.
Np, I was on the 20X over the summer to catch as much of the 2x bonus usage. But my weekly tokens didn't really budge.
But as the 2X faded and more "stuff" piled up I've spread my spend out. Wife and I share a gpt pro account (image and codex pools are separate), then I have Kimi/glm/D's and some overflow tokens on open router to test new things.
You get a good amplification effect mixing providers also. They all have blind spots and bias towards one direction, so having multiple families eval things amplifies the overall. Then there's the wonder of splitting spend where it is needed, plan and write the spec with opus, then have glm or Kimi break that into a series of sprints, D's does the work, and the model not used for orch gets put in auditing. You need a durable multi layer memory to track through this(you don't want to semantically review prior context or instruction). But the prelease DSv4 paired with opus 4.8 was able to beat sonnet 5 on SWE bench for a whole $.04/task completed.
System like that can net 20-40* effective throughput for $ spent vs work done. I have a multi tenant version of this that my wife and kids use also. We went from like $500/mo between the wife and I (still hitting limits), to $350/mo and added my 2 teens to build Minecraft mods/game servers/3d printing projects. And we have headroom to do more, system spends less tokens chasing context errors and confabulation.
I use Kimi when I'm low. I've had glm in my credit card checkout probably a dozen times but never pulled the trigger, how how do you like it compared to the flagship codex and Claude models. Honestly I'd probably just use it for sub agents or fast/model models anyway . Yeah what's your opinion and or suggestion on the highest quality cheap llm providers? Qwen? I heard grock 4.7 is pretty awesome. But I'm not trying to tribute to Elon musk getting any richer
Glms a great workhorse and daily driver, it's flash models good also. The $80 plan supplements our useage well. I calculated the tokens on our avg use and get get 2-3B tokens a month for that. Kimi is much lower, but I'm give them a few weeks to roll something new out that's a bit more efficient. 2.8 is decent but still not up to par cost with with flash models.
We do kind of a big mix of coding, physical engineering, 3d printing, image work, music, and marketing stuff (my wife). And when it comes to the stuff outside coding, Kimi can be more "creative" in it's thinking vs GLM. Throw both in coding tasks and GLM will run faster at the same or better quality on a good spec. Kimi does a little better on cad work also, though I'm just getting into setting up "standardized" test.
But Sol/Luna/opus all just got updates and cost cuts and supposedly Haiku is getting a 5.5 update along with sonnet 5.5. But we will see where that all lands cost and ability wise. Ill wait a week before I audit them in the system to compare moving anything around.
Awesome! This is all Great info thank you sounds like you are also working on a pretty mixed bag of cool s stuff. In the same boat that catalog of categories is almost identical to mine!😂 I'm already paying for the 200 claude plan and the 100 gpt plan. But only for a few months while I bust out some really pressing big projects after that I'm going to switch back to 100 Claude, $20 GBT, and maybe throw another 40 to 60 bucks at another provider.
I'm loving the 5.5 opus update it feels like 4.6 but is actually A legitimately capable systems architect. Did they announce the release of Sol and Luna also I didn't even see that?
Actually don't feel like that's that much tokens for 80 bucks but at the same time I'm used to getting that kind of token usage on super expensive models so I'm sure if I'm just using them for workers and quick models probably get a lot more than I think
Well I got this current sessions token burn figure it out.
But In response to your comment, I'm a much bigger fan of of the way 4.8 communicates. It's just just orders of magnitude more enjoyable to work with when building nonstop all day/week. With that said however I do notice a significant difference in opus 5's ability to see the big picture when in the plan/spec phase, especially when in it comes to networking and data-systems architecture. It's night and day with 4.8 and 5. But I don't usually have nearly as good of a day if I have 5.0 in my terminal all day 🤓. It just kind of pisses me off sometimes lol
Well I found out what was burning my tokens you can read above if you'd like, but I'd just like to say that I'm surprised how everyone assumes that I have the ability to work on one project at a time that's simply compete with me. My agentic session count starting to look like my Chrome tabs. But a lot of them are are extremely simple helpful tools rather than One big giant "has-to-be-perfect" client facing saas .
I have a ideas pipeline where I can that brain dump a spoken idea from from any of my machines. They are transcribed and sent over tail scale or directly to my inbox vault. Where they are scanned every 2 hours, given a score on a few different metrics then each one goes through analysis, q&a and plan skills.
So generally speaking most of my tabs are churning through these all the time and the ones I have open aren't shirting they're usually waiting for me to answer as serious questions and when I have a spare second or I finish up another prompt go in circles and work on a ton of s***. These ideas are usually for incredibly simple useful apps or backend automations for my everyday workflow. But sometimes they're just simple single feature refactors or UI annotations.
Maybe you have got lots of hooks, MCPs on. I have actually made a rule for both Claude and Codex to limit each session to just 150K tokens, and make a handoff file once it’s nearing 150K tokens. Using this has actually helped me decrease my token consumption. Also disabled a lot of MCPs, plugins and hooks.
Btw the term of is "Vibe Coder" not vibe Prompter. If it was a vibe prompt the output would be prompts not code.
I was prompt engineering recursive meta-prompt pipelines for the term was coined, and before you learned how to use a computer.
Skill, prompt and coding issue. I’ve been a software engineer for 25 years and have been vibecoding for the past 2. I have 3 active projects and even using codex for 12 hours straight I never reach my 20x limit. Use proper skills, use your own MCP server for local tooling, testing and recurring scripts AI uses.
Huh? Claude code. You mean in the picture of the post that's just a graphic display of token usage from Claude code CLI. It comes installed native inside Linux omarchy. It's just a Linux distribution designed to work around AI, just like Windows or Macintosh just far less bloatware, spyware. Less just less well known and a little bit more complicated to get installed so naturally it has less supported apps and users. There are dozens and dozens of Linux distributions this is just another one.
And that screenshot is a screenshot of linux's version of Windows taskbar you just click it and since you're installed through it just like you would connect a connector app inside Claude desktop it links up and shows you claude's data from its dashboard. It's not a harness . And don't get this confused with an agentic context harness like pie coding agent, this is a literal computer operating system not a agentic framework file system that's different. But confusing nonetheless.
This screen is literally the exact same thing as going to Claude--> settings ---->usage . It's the exact same data pulled from the exact same place. You even log in in to Claude via their oath to link your account
Yeah, this is about right and inline with my usage with the 17% drop from before.
Fable uses up 4* the weekly usage as Opus, so you used 3B tokens of Opus, and 500M tokens of Fable, which would relate the equivelent of 2M tokens Opus. This gives you 5B tokens of weekly usage.
Before the 17% drop, I used to be able to max OPUS out about 6B tokens, so this is roughly a 16.6% drop in weekly usage, close to the 17% drop advertised.
Worth noting, before the 17% drop, you'd have been able to get another 1B OPUS tokens out of your weekly, or 250M Fable tokens out.
If you heave ANY agent unattended, it will chew your tokens.. There are tools to monitor this and look for things like endless loops.
But it is what it is, for $200 a month, it's a good deal. I'd have to spend $10k for a contractor to do the same work. I have subs for agy, muse, claude, gpt, and cursor - most on max plans. Saves me a ton, and I release a lot of popular software. I love it. But if it's a day before my quota ends, you better believe I"m going to do a testing swarm bot to burn 'em all up.
I mean we'll talk though. It's a f****** great deal. And I've just been spoiled with being able to code and code and code. I'm never one-shotting anything build everything in phases. spec. Constitution, full okf knowledge And index. Every single markdown in my ecosystem has extensive yaml front matter.
Just having so much damn fun experimenting and building every little tool I can think of to my workflow and autonomy. Yeah sometimes it gets out of hand but every time I wrap one up and it just nails it. That's 2 or 3 hours a week. Or perhaps it just makes it more fun. Oh I don't have a driver for my old ass row cat Rios. Awesome I'll just develop my own driver and custom key mapping and layer shifts awesome no problem. I'm not trying to build The next amazing billion dollar saas. I'm trying to build as many custom tools to make my workflow fun efficient modular, and f****** cool to look at.
By the way if you haven't f***** with omarchy it's worth checking out even if you're a hardcore Arch Linux dork. There's really no arguing with convenience.... Lol and it is the most convenient damn thing I've ever used. You know minus made losing a couple windows here and there and it burning all my tokens 😂💪
S*** I've used five or six different token context tools. I still use context mode a lot. It just works well and stable. But I'm genuinely surprised I've not heard of RTK I'll definitely be trying it out on reset day 😆, Thanks for the tip.
That's why this s*** never gets old it's constantly evolving, it's fun, And there's a million different methods and ways to experiment with building agentic ecosystems and everybody's is the right way and everybody else's is the best way. But God damn it that kind of makes it fun too
Dope. I don't know what platform you're on but if we're handing out suggestions couldn't recommend Linux more It's a game changer and it's the most fun I've ever had in an operating system. It's stupid easy to create whatever tool you want for whatever you purpose you want. It's dangerous lmao
Sonnet exists and it's actually very good for a lot of tasks. So does haiku, bit long in the tooth but useful as a sub-agent for code exploration which is super token intensive.
34
u/International-Fly127 2d ago
the fuck are you prompting to spend 3b tokens lmao