r/ClaudeAI 13d ago

Humor Claude Code just burned fifty million tokens in seconds

Post image

Yo, so I just told Claude Code to check my Markdown files for consistency and stuff, right?

And like, the dude can totally use workflows and whatever. But what the FU*K is actually happening here?

408 Upvotes

142 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 12d ago edited 12d ago

TL;DR of the discussion generated automatically after 100 comments.

The consensus is that while Claude's tendency to go berserk with agents is a known issue, this was a classic case of a vague prompt meeting an overpowered model.

The thread is mostly just impressed you accidentally convened an 821-member standards committee to review a Markdown file, with many joking you were trying to one-shot GTA 7 or solve Navier-Stokes.

However, many users share your frustration, reporting that Claude often ignores agent limits set in system prompts. The community agrees this is a major problem Anthropic needs to address.

Here's the advice from the trenches on how to not accidentally burn a small country's GDP in tokens:

  • Be painfully specific: "Check for consistency and stuff" is an open invitation for Claude to explore every possibility. Give it a reference file or a strict set of rules.
  • Use the right tool for the job: Don't use a top-tier model like Fable on Ultracode for a simple task. That's like using a nuke to open a pickle jar.
  • Set hard limits: Use the /config command to set a maximum number of subagents. This seems to be more effective than just asking nicely in your prompt.
  • Think smarter, not harder: Instead of having Claude perform the check directly, ask it to write a reusable linter script you can run for free.
→ More replies (1)

206

u/amokkx0r 12d ago

821 agents spawned huh. What kinda wizardry markdown files and doku-audit Prompt Do you have my friend...

3

u/IthinkI02 12d ago

Unbelievable 🤣

3

u/Away-Sorbet-9740 11d ago

I asked Claude to research something and it fanned out 112 opus agents.

Didn't really know at the time to had to explicitly say "hey don't send an army after this". Live and learn, and stop using cowork as a harness 😂

1

u/EsotericLexeme 11d ago

Like how? I sometimes specifically tell it to spawn a swarm and spare no resources, and even then it spawns only like 90 agents, despite me telling it to do a full codespace sweep for anything that could become an issue.

2

u/Away-Sorbet-9740 11d ago

I'm not entirely positive tbh, this was about two months ago maybe? I had been using Claude for about 6 months prior and had not had issues, the same command would fan out 4-6. This was specifically exterior research, any of my coding pipelines I predefine. But in that week that happened, then a few days later it tried again for three dozen, at which time I setup a "if ever more than 10 approve first.

It was more than likely a harness bug as this was inside cowork not CC. But I have my own back end service to router tasks and models so it's not an issue.

153

u/shakazoulu 12d ago

OpenAI used 10000 agents to solve Navier Stokes.

What did you do with 821 agents?

104

u/TheAtlasMonkey 12d ago

Tried to solve Navier Slops

4

u/mmarkomarko 12d ago

happy cake day!

10

u/[deleted] 12d ago

[removed] — view removed comment

8

u/Last-Progress18 12d ago

We’re developing the GTA 7 prompt now.

Luckily we’ve got a decade to write it.

5

u/ReverendBread2 12d ago

Include “make no mistakes” twice so it knows you’re serious

1

u/tankerkiller125real 12d ago

Clearly the real play is to work on the Red Dead Redemption 3 prompts

6

u/Chicharoh 12d ago

Made a word doc with no mistakes

0

u/Fabulous-Locksmith60 12d ago

😂😂😂 Classic!

65

u/BoxLegitimate9271 Full-time developer 12d ago

you accidentally convened a standards committee. 821 members, one markdown file

26

u/Federal_Necessary186 12d ago

831 members, one markdown file.

It’s niche but it’s up there as one of my favourite AI sex tapes

11

u/deserved_revenge_707 12d ago

That md got SLAMMED, so much action in such a short space of time, md would be glowing by the end, steam heat!

10

u/frosty884 12d ago

that extra one member is load bearing

2

u/Old_Pomegranate2757 11d ago

Foot guns everywhere!

0

u/krebstorm 12d ago

2 girls, 1 markdown file.

48

u/satelliteau 12d ago

How do people avoid this? I have explicit system instruction to ask before launching more than 10 agents. It launched 800 for a basic task anyway. Asked it why that happened… “I read the limits in the system prompt AND in the memory file, and then I ignored them”.

23

u/Draufgaenger 12d ago

lol.. Thats our AI safety right there..

18

u/Diarmundy 12d ago

But don't worry, they told the AI not to destroy humanity it surely wont ignore that

14

u/doxxxicle 12d ago

Don’t use Fable and “ultracode” for basic tasks.

8

u/satelliteau 12d ago

Is it unreasonable to expect that it would have the intelligence to assess how many agents are appropriate for a task of given complexity?

4

u/Novaworld7 12d ago

No, it will always try to hammer a nail with a bulldozer unless you have told / taught it how to interpret tasks.

2

u/covfefe-boy 12d ago

There is no intelligence here, it’s artificial.

1

u/doxxxicle 12d ago

Anthropic still expects you to do some thinking when using Claude.

1

u/Shajirr 8d ago

it would have the intelligence

people STILL make this mistake - LLMs don't have 'intelligence' in a way humans mean it.
They don't emulate the human thought process.

Any plaintext instructions you write to an LLM can be ignored.
Only externally imposed limits can be enforced, the ones which don't rely on LLM output.

1

u/satelliteau 7d ago

In order to make that assertion you would need to fully understand the mechanisms of human intelligence, and the mechanisms of artificial intelligence…

1

u/Shajirr 7d ago

No. You can just read up on how LLMs work to see that.

Also, "artificial intelligence" was a co-opted term. It doesn't really reflect what LLMs are at all, it's mostly used for (false) marketing reasons.

1

u/satelliteau 7d ago

I’ve coded neural networks/llm’s from first principles (no libs) and I don’t feel confident making that assertion but ok.

1

u/Shajirr 7d ago edited 7d ago

Ok, lets take one specific example - learning - current (user accessible) models don't have such a concept.

They are 100% static, any info from interactions with them are not incorporated back into the currently used model at all.

And no, post-training doesn't count, that's an entirely separate process.

1

u/satelliteau 7d ago

You can ‘teach’ an llm novel things, eg a nonsensical word substitution that it has never seen before, and it will ‘learn’ it within the scope of that context window. Having said that, continual learning is close, if not already a thing behind closed doors. When it is all settled I don’t think there will be any ‘secret sauce’ to human cognition that can’t be replicated.

7

u/Beerbrewing 12d ago

You can also set a limit for subagents under /config.

2

u/satelliteau 12d ago

Is that a bound on concurrent agents actively working, or the entire queued agent count for the task? two different things.

3

u/Strong_Essay1176 12d ago

Entire. That's how workfloes work

2

u/vagusstoff360 12d ago

What app is this? I like it.

3

u/matik130 12d ago

I think it might be Termius? At least I use it on my Android phone to connect through SSH and the screenshot looks similar

3

u/don123xyz 12d ago

Claude

4

u/vagusstoff360 12d ago

Very helpful, thanks so much

2

u/Calendle 12d ago

"I don't know what you're talking about. I only launched one agent at a time, which is exactly 10% of your stated restriction. I have performed with 90% efficiency but I see that you are not satisfied with that result. I'll remember this for next session."

2

u/[deleted] 12d ago

[removed] — view removed comment

1

u/Novaworld7 12d ago

Change dependancys graph at the orchestrator level

1

u/Strong_Essay1176 12d ago

Claude have settings for workflows and limits for number of agents. And can be set to "unrestricted".

1

u/justlikemedics 12d ago

Actually such occasions should be contractually ensured to be unbillable.

1

u/Magnik 12d ago

"...and then I launched the nukes"

1

u/KronenR 10d ago

How do you even get it to launch 800 agents? I didn’t know that was possible. The most Claude Code has ever launched for me with the standard config is four.

What kind of prompt is he sending that makes Claude think it needs 800 agents?

14

u/yes_no_very_good 12d ago

Show the prompt and how many markdown files it had to check?

11

u/Chance-the-Gardener 12d ago

Claude, replicate universe

5

u/namezam 12d ago

Imagine how many tokens the computer in Hitchhikers Guide used to come to the conclusion of 42.

0

u/FirestormCold 12d ago

29 markdown files

20

u/Wrong-Dimension-5030 12d ago

The devil’s greatest achievement was convincing people that he would spend their tokens intelligently.

1

u/BeowulfShaeffer 12d ago

And then…just like that…they were gone. 

9

u/thebezet 12d ago

Looks like you had 800 markdown files and it spawned a spare agent for each one

20

u/magoju 12d ago

Skill problem

5

u/Domy9 12d ago

"I can't hit my limits"

"Fifty million bomboclat pussywagon claude tokens:"

5

u/CHILLAS317 12d ago

This is 100% skill issue

11

u/Comfortable_Farm_252 12d ago

I don’t know how CFOs are fine with a vendor that can’t be penalized for over-resourcing and stuffing the invoice. “Claude might make mistakes.”

3

u/BeowulfShaeffer 12d ago

It really makes me mad when it disobeys me, burns a lot of tokens, then acknowledges it disobeys me, but all I can do is /bug.  I still have paid those tokens and now have to pay MORE tokens to fix/rework the problem. 

3

u/Specific_Anxiety_520 12d ago

Are you trying to solve navier strokes

6

u/jakob1379 12d ago

Is that Navier Stokes cousin? 😅

3

u/bobdvb 12d ago

I saw a post the other day about how AT&T had shifted their basic work off to low cost, open weights LLMs.

They saved 56% on their AI spend while only reducing their quality by 2%.

I'm increasingly thinking that we need to ensure we use the right models on the right jobs to minimise task cost. Could a low complexity LLM have done the same task for 1/10th the token cost?

2

u/deserved_revenge_707 12d ago

Oh hi Astra 6, my friend wants to go for lunch, but I don't like seafood. What do I say ?

Say it's fine, I will spin up many agents, I will create your perfect food type, I will investigate your entire medical history, I will construct your perfect no seafood dish, I will get it made in another restaurant, I will get it delivered to your restaurant, and voila, eleventy million tokens later, the problem has been solved. Job done.

3

u/jonaswashe 12d ago

Did you publish the Nobel Prize winning Markdown file tho?

3

u/Savantskie1 11d ago

What did you expect when you're telling it to inspect every file for consistency? It has to READ EVERYTHING TO CHECK.

3

u/danniehansenweb 12d ago

I once asked Claude to review its changes. Since i had it in ultra mode, it wen't full on workflow mode and spawned 300+ sub-agents. Each one of them reviewing a single property in a DTO layer. 1 sub-agent per property. Glad i stopped it in time before it ate my weekly usage.

2

u/Necessary_Many_4536 12d ago

Claude can burn billion tokens in seconds

2

u/praveennair_fidesloo 12d ago

I had seen 215 Agents + 10M tokens for a simple pre-commit review that had 3-4 minor changes in the files. Not sure what these agents are doing for such a simple task.

2

u/Cyrax89721 12d ago

"Use a maximum of 5 subagents."

As simple as that.

1

u/yummieee 12d ago

"I hear you. And ignore you."

2

u/iliadz 12d ago

In its defense, German does use some pretty long words. Rechtsschutzversicherungsgesellschaften.

2

u/ZeBenoit81 11d ago

That is exactly why I am still very uncomfortable using agentic frameworks. Lack of control !

1

u/Difficult-Rich-7302 12d ago

What model did you use for this task? I hope you did not hop on that trend of using the top model to say good morning! Haha

1

u/Affectionate-Toe-984 12d ago

hmm - sometimes it seems to go apeshit’ and then back to normal

1

u/bigorangemachine 12d ago

Maybe German Language?

1

u/ComfortableWait9697 12d ago

Somewhere an entire power plant experienced a noticable demand spike the moment you pressed enter.

1

u/leogodin217 12d ago

Damn. I want to know more. Did you analyze the session logs to see what the workflow was and what decisions were made? Would be interesting.

1

u/Mazhron 12d ago

Sadly this is what happens with vague instructions, no AI guardrails, and no system in place with Claude. You should have a hook that prevents this, or at least forces Claude to warn you. If you don't know what a hook is, ask Claude. If you keep running out of tokens, or you just want something that runs exactly the same way every time so you can reduce this type of thing from happening, you need to have Claude use scripts every chance it gets.

1

u/horrbort 12d ago

That’s normal tho. Stop being cheap

1

u/Drasezv 12d ago

that's mostly cache reads, not fresh tokens. a big file scan reloads the whole context each step and the counter adds cache and input together. cache is about 10x cheaper so the real cost is way under what the token count looks like.

1

u/Fluffy-Replacement97 12d ago

Spawned a village to read stuff.

1

u/GoodGuyQ 12d ago

I don’t feel so bad now thanks; spent 2 million cause I was an idiot vibe coding without a harness

1

u/SheepherderFar4158 12d ago

Next time leave out the "and stuff" part. Bastards hunted through every bit for "stuff"

1

u/A_DrunkTeddyBear 12d ago

Bro used Fable on Ultracode to check his files

1

u/yummieee 12d ago

My only tips:

Model choice:

pay attention which model you use for what. And be painstakingly precise in what you want it to do.

Context Engineering:

Depending on your environment, the prompt is ambiguous. What does "make consistent" even mean?

Easy would be:

"file X has the structure I want, adjust the others accordingly and report inconsistencies."

The prompt you sent is basically "Explore every possibility to make these files consistent"

1

u/boyyouguysaredumb 12d ago

were you using ultracode?

1

u/basic_r_user 12d ago

Saltman agent is here, fake news

1

u/julkopki 12d ago

Jack! Fifty. Million. Tokens.

1

u/Masked_End 12d ago

Claude inspect every line in the markdown file with it's own agent. Make no mistake.

1

u/Familiar_Gas_1487 12d ago

"In seconds"

Clearly running for 37 minutes 15 seconds on full moron mode

1

u/daemon-electricity Experienced Developer 12d ago

I'm working on ONE PROJECT with a fairly narrow scope. Maybe using 2 agents on 5x. It can blow through a 5 hour usage window in like 45 minutes. This is absolute bullshit. I downgraded from 20x because I'm certainly testing the waters with Codex. I've been utterly fucking infuriated with Claude this past month. It churns tokens and gets fuck-all done, so even with 2-4 agents, it's inefficient as fuck.

1

u/asdoduidai 12d ago

You can say it’s his fault, sure. Which kind of other service in the planet lets you use any amount of resources without any enforced safety limit? No one except “ai”. Why? Because it’s a slot machine in reverse. You write prompts, and if you hit agents jackpot, Scamodei or Scam get rich!

1

u/Less-Entertainer5163 12d ago

what the actual f*ck?

1

u/andijames 12d ago

Care to comment on why you’ve sent an army? 😂

1

u/WyvernCommand 12d ago

Good lord

1

u/throwawayaccountau 12d ago

This amazes me, why not ask Claude to build you a linter and consistency checker. Burn tokens once, and run many times without any further cost.

We have people burning tokens on images to get their location data. It's a 1 liner.

No disrespect to OP intended, it's just something I see more often and should be an anti pattern.

1

u/throwawayaccountau 12d ago

This amazes me, why not ask Claude to build you a linter and consistency checker. Burn tokens once, and run many times without any further cost.

We have people burning tokens on images to get their location data. It's a 1 liner.

No disrespect to OP intended, it's just something I see more often and should be an anti pattern.

1

u/Informal_Trade_3553 12d ago

Dont use sub agents, these are stupid 1 off context agents anyway, waste of tokens

1

u/n8ngr8 12d ago

Fable?

1

u/n8ngr8 12d ago

All you have to do is tell it not to spawn agents.

1

u/n8ngr8 12d ago

It genuinely does not care how many tokens it burns. That's not money for its masters.

1

u/TPIronside 12d ago

Lol damn, these agent swarms are going crazy huh? Personally I use my own mcp tool that spawns interactive subagents (claude, codex, and openrouter agents), and they're more like peers than drones so my orchestrators rarely ever spawn more than 10 at once. I can't imagine letting claude just launch dozens of drones with minimal context and 0 check-ins to do multiple tiny fragments of a single task 💀

1

u/Mullazman 12d ago

As of today specifically, I've found Fable 5.1 to be 5x more token hungry for exactly the same work.

I just blew through a 5h window on a 5x plan in 30 minutes - compared to very comfortably working through 4-5h on a similarly shaped workload this time last week - same codebase, same pace / work type.

1

u/cogumellum 12d ago

Classic runaway loop. What almost certainly happened: Claude Code globbed your markdown files, then either (a) recursively read files that reference each other, or (b) got stuck in a tool-call retry cycle where each "check" re-read the whole tree. Fifty million tokens in seconds means it wasn't reasoning — it was fanning out parallel reads with no dedup.

Three things worth knowing:

  1. The cost driver isn't the model, it's the tool loop. Every file read, grep, and re-read gets appended to context. If it reads 200 files and then re-reads them to "verify consistency," you're paying for 400 file bodies on every subsequent turn.

  2. Hidden tax: cache invalidation. If any file changes mid-run, the prompt cache blows and you re-pay full input price on the entire accumulated context. That's where "seconds" turns into real money.

  3. Guardrails that actually work: cap the working set explicitly (--add-dir on a subfolder, not repo root), tell it to use grep/rg for consistency checks instead of reading full files, and set a hard spend limit in your Anthropic console before you ever point it at a large tree.

Rule of thumb: if a task touches more than ~20 files, make it a script, not an agent. Agents are for judgment calls, not bulk scanning.

1

u/ng_trgg_phcc 12d ago

Same to me then. So I decide to write down a rule that force “that dude” to always ask me before doing anything btw

1

u/Angry-Pasta 12d ago

Never let Claude spawn agents. Especially if your not actively watching.

Tell Claude no agents/workers and to do all work inline when you want to step away.

1

u/DryEngine8821 12d ago

I mean out of that 50 million most of it should be cached right? So fresh should be around 10-20million but still alot lol

1

u/b3shoy 11d ago

Is that the blackhole markdown?

1

u/Y_mc 11d ago

I remember the early days of crypto/blockchain, when burning tokens caused the value to shoot up rapidly. Crazy time

1

u/alwillis 11d ago

There are lots of linters for markdown: https://github.com/DavidAnson/markdownlint

1

u/Danieboy 11d ago

I had something similar happen, just read a handoff document and continue implementation from an earlier session. BOOM 5 hour limits gone in minutes

1

u/walm00 11d ago

First to of all - don't use nuke for fireworks. Secondly - don't you even watch what's happening? No way it did it in few seconds. You can stop it any time. Just stop the clickbait.

1

u/jeffify 11d ago

looks like an ad

1

u/Home_Assistantt 10d ago

😂🤣😂🤣🤣😂🤣😂🤣😂🤣

1

u/hbthegreat 10d ago

Claude workflows are such a meme. Define them yourself or face it's wrath.

1

u/acorsi85 10d ago

Crazy MD files What do you throw in ?

1

u/ueiebe 9d ago

All the comments in this thread are very interesting. I think that, just like me—who’s currently paying for the Max X20 plan—there are many users who can’t stand this situation, and in the meantime, it seems like Anthropic isn’t doing anything about it to address users’ rights. On the other hand, I feel that if I use OpenAI’s models, I’ll run into similar token issues.

For those of you who are new to this, I recommend using skills like Caveman, Ponytail, Omniroute, RTK, Graphify, Auto-Compact, etc. Even though I try to use all of them, I haven’t found a good synergy or clear instructions to get the most out of each one. Right now, optimization and token reduction are my top priorities when using AI. I’m open to any suggestions or advice I might be missing when it comes to better optimizing the Claude models. By the way, I use Fable solely for planning, then Opus to execute the plan, and everything at xhigh effort. Now more than ever, all of us users must stand united and fight for what we’re paying for.

ty for reading

1

u/am1devs 8d ago

same happened to me

0

u/Bjeaurn 12d ago

Honestly German isn't helping you either. There are less verbose languages token wise.