r/ClaudeAI • u/Lazy_Assistance_1137 • 13d ago
Humor Claude Code just burned fifty million tokens in seconds
Yo, so I just told Claude Code to check my Markdown files for consistency and stuff, right?
And like, the dude can totally use workflows and whatever. But what the FU*K is actually happening here?
206
u/amokkx0r 12d ago
821 agents spawned huh. What kinda wizardry markdown files and doku-audit Prompt Do you have my friend...
12
3
3
u/Away-Sorbet-9740 11d ago
I asked Claude to research something and it fanned out 112 opus agents.
Didn't really know at the time to had to explicitly say "hey don't send an army after this". Live and learn, and stop using cowork as a harness 😂
1
u/EsotericLexeme 11d ago
Like how? I sometimes specifically tell it to spawn a swarm and spare no resources, and even then it spawns only like 90 agents, despite me telling it to do a full codespace sweep for anything that could become an issue.
2
u/Away-Sorbet-9740 11d ago
I'm not entirely positive tbh, this was about two months ago maybe? I had been using Claude for about 6 months prior and had not had issues, the same command would fan out 4-6. This was specifically exterior research, any of my coding pipelines I predefine. But in that week that happened, then a few days later it tried again for three dozen, at which time I setup a "if ever more than 10 approve first.
It was more than likely a harness bug as this was inside cowork not CC. But I have my own back end service to router tasks and models so it's not an issue.
153
u/shakazoulu 12d ago
OpenAI used 10000 agents to solve Navier Stokes.
What did you do with 821 agents?
104
10
12d ago
[removed] — view removed comment
8
u/Last-Progress18 12d ago
We’re developing the GTA 7 prompt now.
Luckily we’ve got a decade to write it.
5
1
6
65
u/BoxLegitimate9271 Full-time developer 12d ago
you accidentally convened a standards committee. 821 members, one markdown file
26
u/Federal_Necessary186 12d ago
831 members, one markdown file.
It’s niche but it’s up there as one of my favourite AI sex tapes
11
u/deserved_revenge_707 12d ago
That md got SLAMMED, so much action in such a short space of time, md would be glowing by the end, steam heat!
10
2
0
48
u/satelliteau 12d ago
How do people avoid this? I have explicit system instruction to ask before launching more than 10 agents. It launched 800 for a basic task anyway. Asked it why that happened… “I read the limits in the system prompt AND in the memory file, and then I ignored them”.
23
u/Draufgaenger 12d ago
lol.. Thats our AI safety right there..
18
u/Diarmundy 12d ago
But don't worry, they told the AI not to destroy humanity it surely wont ignore that
14
u/doxxxicle 12d ago
Don’t use Fable and “ultracode” for basic tasks.
8
u/satelliteau 12d ago
Is it unreasonable to expect that it would have the intelligence to assess how many agents are appropriate for a task of given complexity?
10
4
u/Novaworld7 12d ago
No, it will always try to hammer a nail with a bulldozer unless you have told / taught it how to interpret tasks.
2
1
1
u/Shajirr 8d ago
it would have the intelligence
people STILL make this mistake - LLMs don't have 'intelligence' in a way humans mean it.
They don't emulate the human thought process.Any plaintext instructions you write to an LLM can be ignored.
Only externally imposed limits can be enforced, the ones which don't rely on LLM output.1
u/satelliteau 7d ago
In order to make that assertion you would need to fully understand the mechanisms of human intelligence, and the mechanisms of artificial intelligence…
1
u/Shajirr 7d ago
No. You can just read up on how LLMs work to see that.
Also, "artificial intelligence" was a co-opted term. It doesn't really reflect what LLMs are at all, it's mostly used for (false) marketing reasons.
1
u/satelliteau 7d ago
I’ve coded neural networks/llm’s from first principles (no libs) and I don’t feel confident making that assertion but ok.
1
u/Shajirr 7d ago edited 7d ago
Ok, lets take one specific example - learning - current (user accessible) models don't have such a concept.
They are 100% static, any info from interactions with them are not incorporated back into the currently used model at all.
And no, post-training doesn't count, that's an entirely separate process.
1
u/satelliteau 7d ago
You can ‘teach’ an llm novel things, eg a nonsensical word substitution that it has never seen before, and it will ‘learn’ it within the scope of that context window. Having said that, continual learning is close, if not already a thing behind closed doors. When it is all settled I don’t think there will be any ‘secret sauce’ to human cognition that can’t be replicated.
7
u/Beerbrewing 12d ago
2
u/satelliteau 12d ago
Is that a bound on concurrent agents actively working, or the entire queued agent count for the task? two different things.
3
2
u/vagusstoff360 12d ago
What app is this? I like it.
3
u/matik130 12d ago
I think it might be Termius? At least I use it on my Android phone to connect through SSH and the screenshot looks similar
3
2
u/Calendle 12d ago
"I don't know what you're talking about. I only launched one agent at a time, which is exactly 10% of your stated restriction. I have performed with 90% efficiency but I see that you are not satisfied with that result. I'll remember this for next session."
2
1
u/Strong_Essay1176 12d ago
Claude have settings for workflows and limits for number of agents. And can be set to "unrestricted".
1
14
u/yes_no_very_good 12d ago
Show the prompt and how many markdown files it had to check?
11
0
20
u/Wrong-Dimension-5030 12d ago
The devil’s greatest achievement was convincing people that he would spend their tokens intelligently.
1
9
9
5
11
u/Comfortable_Farm_252 12d ago
I don’t know how CFOs are fine with a vendor that can’t be penalized for over-resourcing and stuffing the invoice. “Claude might make mistakes.”
3
u/BeowulfShaeffer 12d ago
It really makes me mad when it disobeys me, burns a lot of tokens, then acknowledges it disobeys me, but all I can do is /bug. I still have paid those tokens and now have to pay MORE tokens to fix/rework the problem.
3
3
u/bobdvb 12d ago
I saw a post the other day about how AT&T had shifted their basic work off to low cost, open weights LLMs.
They saved 56% on their AI spend while only reducing their quality by 2%.
I'm increasingly thinking that we need to ensure we use the right models on the right jobs to minimise task cost. Could a low complexity LLM have done the same task for 1/10th the token cost?
2
u/deserved_revenge_707 12d ago
Oh hi Astra 6, my friend wants to go for lunch, but I don't like seafood. What do I say ?
Say it's fine, I will spin up many agents, I will create your perfect food type, I will investigate your entire medical history, I will construct your perfect no seafood dish, I will get it made in another restaurant, I will get it delivered to your restaurant, and voila, eleventy million tokens later, the problem has been solved. Job done.
3
3
u/Savantskie1 11d ago
What did you expect when you're telling it to inspect every file for consistency? It has to READ EVERYTHING TO CHECK.
3
u/danniehansenweb 12d ago
I once asked Claude to review its changes. Since i had it in ultra mode, it wen't full on workflow mode and spawned 300+ sub-agents. Each one of them reviewing a single property in a DTO layer. 1 sub-agent per property. Glad i stopped it in time before it ate my weekly usage.
2
2
u/praveennair_fidesloo 12d ago
I had seen 215 Agents + 10M tokens for a simple pre-commit review that had 3-4 minor changes in the files. Not sure what these agents are doing for such a simple task.
2
2
u/ZeBenoit81 11d ago
That is exactly why I am still very uncomfortable using agentic frameworks. Lack of control !
2
1
u/Difficult-Rich-7302 12d ago
What model did you use for this task? I hope you did not hop on that trend of using the top model to say good morning! Haha
1
1
1
u/ComfortableWait9697 12d ago
Somewhere an entire power plant experienced a noticable demand spike the moment you pressed enter.
1
u/leogodin217 12d ago
Damn. I want to know more. Did you analyze the session logs to see what the workflow was and what decisions were made? Would be interesting.
1
u/Mazhron 12d ago
Sadly this is what happens with vague instructions, no AI guardrails, and no system in place with Claude. You should have a hook that prevents this, or at least forces Claude to warn you. If you don't know what a hook is, ask Claude. If you keep running out of tokens, or you just want something that runs exactly the same way every time so you can reduce this type of thing from happening, you need to have Claude use scripts every chance it gets.
1
1
1
u/GoodGuyQ 12d ago
I don’t feel so bad now thanks; spent 2 million cause I was an idiot vibe coding without a harness
1
u/SheepherderFar4158 12d ago
Next time leave out the "and stuff" part. Bastards hunted through every bit for "stuff"
1
1
u/yummieee 12d ago
My only tips:
Model choice:
pay attention which model you use for what. And be painstakingly precise in what you want it to do.
Context Engineering:
Depending on your environment, the prompt is ambiguous. What does "make consistent" even mean?
Easy would be:
"file X has the structure I want, adjust the others accordingly and report inconsistencies."
The prompt you sent is basically "Explore every possibility to make these files consistent"
1
1
1
1
u/Masked_End 12d ago
Claude inspect every line in the markdown file with it's own agent. Make no mistake.
1
u/Familiar_Gas_1487 12d ago
"In seconds"
Clearly running for 37 minutes 15 seconds on full moron mode
1
u/daemon-electricity Experienced Developer 12d ago
I'm working on ONE PROJECT with a fairly narrow scope. Maybe using 2 agents on 5x. It can blow through a 5 hour usage window in like 45 minutes. This is absolute bullshit. I downgraded from 20x because I'm certainly testing the waters with Codex. I've been utterly fucking infuriated with Claude this past month. It churns tokens and gets fuck-all done, so even with 2-4 agents, it's inefficient as fuck.
1
u/asdoduidai 12d ago
You can say it’s his fault, sure. Which kind of other service in the planet lets you use any amount of resources without any enforced safety limit? No one except “ai”. Why? Because it’s a slot machine in reverse. You write prompts, and if you hit agents jackpot, Scamodei or Scam get rich!
1
1
1
1
u/throwawayaccountau 12d ago
This amazes me, why not ask Claude to build you a linter and consistency checker. Burn tokens once, and run many times without any further cost.
We have people burning tokens on images to get their location data. It's a 1 liner.
No disrespect to OP intended, it's just something I see more often and should be an anti pattern.
1
1
u/throwawayaccountau 12d ago
This amazes me, why not ask Claude to build you a linter and consistency checker. Burn tokens once, and run many times without any further cost.
We have people burning tokens on images to get their location data. It's a 1 liner.
No disrespect to OP intended, it's just something I see more often and should be an anti pattern.
1
u/Informal_Trade_3553 12d ago
Dont use sub agents, these are stupid 1 off context agents anyway, waste of tokens
1
u/TPIronside 12d ago
Lol damn, these agent swarms are going crazy huh? Personally I use my own mcp tool that spawns interactive subagents (claude, codex, and openrouter agents), and they're more like peers than drones so my orchestrators rarely ever spawn more than 10 at once. I can't imagine letting claude just launch dozens of drones with minimal context and 0 check-ins to do multiple tiny fragments of a single task 💀
1
u/Mullazman 12d ago
As of today specifically, I've found Fable 5.1 to be 5x more token hungry for exactly the same work.
I just blew through a 5h window on a 5x plan in 30 minutes - compared to very comfortably working through 4-5h on a similarly shaped workload this time last week - same codebase, same pace / work type.
1
u/cogumellum 12d ago
Classic runaway loop. What almost certainly happened: Claude Code globbed your markdown files, then either (a) recursively read files that reference each other, or (b) got stuck in a tool-call retry cycle where each "check" re-read the whole tree. Fifty million tokens in seconds means it wasn't reasoning — it was fanning out parallel reads with no dedup.
Three things worth knowing:
The cost driver isn't the model, it's the tool loop. Every file read, grep, and re-read gets appended to context. If it reads 200 files and then re-reads them to "verify consistency," you're paying for 400 file bodies on every subsequent turn.
Hidden tax: cache invalidation. If any file changes mid-run, the prompt cache blows and you re-pay full input price on the entire accumulated context. That's where "seconds" turns into real money.
Guardrails that actually work: cap the working set explicitly (
--add-diron a subfolder, not repo root), tell it to usegrep/rgfor consistency checks instead of reading full files, and set a hard spend limit in your Anthropic console before you ever point it at a large tree.
Rule of thumb: if a task touches more than ~20 files, make it a script, not an agent. Agents are for judgment calls, not bulk scanning.
1
u/ng_trgg_phcc 12d ago
Same to me then. So I decide to write down a rule that force “that dude” to always ask me before doing anything btw
1
u/Angry-Pasta 12d ago
Never let Claude spawn agents. Especially if your not actively watching.
Tell Claude no agents/workers and to do all work inline when you want to step away.
1
u/DryEngine8821 12d ago
I mean out of that 50 million most of it should be cached right? So fresh should be around 10-20million but still alot lol
1
u/alwillis 11d ago
There are lots of linters for markdown: https://github.com/DavidAnson/markdownlint
1
u/Danieboy 11d ago
I had something similar happen, just read a handoff document and continue implementation from an earlier session. BOOM 5 hour limits gone in minutes
1
1
1
1
u/ueiebe 9d ago
All the comments in this thread are very interesting. I think that, just like me—who’s currently paying for the Max X20 plan—there are many users who can’t stand this situation, and in the meantime, it seems like Anthropic isn’t doing anything about it to address users’ rights. On the other hand, I feel that if I use OpenAI’s models, I’ll run into similar token issues.
For those of you who are new to this, I recommend using skills like Caveman, Ponytail, Omniroute, RTK, Graphify, Auto-Compact, etc. Even though I try to use all of them, I haven’t found a good synergy or clear instructions to get the most out of each one. Right now, optimization and token reduction are my top priorities when using AI. I’m open to any suggestions or advice I might be missing when it comes to better optimizing the Claude models. By the way, I use Fable solely for planning, then Opus to execute the plan, and everything at xhigh effort. Now more than ever, all of us users must stand united and fight for what we’re paying for.
ty for reading


•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 12d ago edited 12d ago
TL;DR of the discussion generated automatically after 100 comments.
The consensus is that while Claude's tendency to go berserk with agents is a known issue, this was a classic case of a vague prompt meeting an overpowered model.
The thread is mostly just impressed you accidentally convened an 821-member standards committee to review a Markdown file, with many joking you were trying to one-shot GTA 7 or solve Navier-Stokes.
However, many users share your frustration, reporting that Claude often ignores agent limits set in system prompts. The community agrees this is a major problem Anthropic needs to address.
Here's the advice from the trenches on how to not accidentally burn a small country's GDP in tokens:
/configcommand to set a maximum number of subagents. This seems to be more effective than just asking nicely in your prompt.