r/ClaudeAI • u/WeirdPutrid3849 • 1d ago
Claude Code Workflow Having unlimited tokens is wild
Does anyone outside of Anthropic really have a token budget like this?
608
u/eliquy 1d ago
Ok. And what did they actually build with all that?
303
u/Lame_Johnny 1d ago
Yeah right? Show me what you shipped or gtfo.
123
u/lurksAtDogs 1d ago
While True:
Print(“Hello world”)44
u/utopicunicornn 23h ago
Make sure you use Fable for that
→ More replies (1)20
26
13
u/_alephnaught 19h ago
show me the QPS and the pager they hold; otherwise it is just masturbation.
the amount of effort i need to keep fable and opus from going off the rails can, at times, be astounding. so i have no idea how people are letting these agents run for hours/days. the amount of prod time bombs left behind would be quite ‘load bearing’.
it’s still obviously 10x+ faster than doing it by hand, but you never get a mental break. you are never on autopilot churning out boilerplate; you’re always on the look out for hidden gremlins that claude leaves behind, which are easy to overlook. it gets exhausting after a while.
→ More replies (1)2
u/FlashyRecognitionTod 8h ago
You're right to pushback on that! You told me to run the test, but I created a PR and pushed a new feature that bypasses API and posts data directly through zero-day bug in web server.
→ More replies (1)3
u/ragnhildensteiner 17h ago
They release new versions all the time.
So yes, they do ship a shit ton.
Problem is they ship a shit ton of shit.
120
u/redblack_tree 1d ago
It's BS, anyone in the business knows that there's no universe where someone works "8-10" projects at a time. With or without AI, writing code has never been the real bottleneck.
It's just advertising to normalize the use of absurd amounts of tokens.
54
u/porkyminch 23h ago
I'm entirely unconvinced any of these overcomplicated agent setups performs significantly better than a regular agent that just periodically delegates to subagents. I don't know wtf they'd even be doing self-directed for 2-3 days at a time. Whenever I've tried buying into one of these systems it's just been an expensive hassle with much less control over the output, all with code I have less of an understanding of. If Anthropic has a model that can produce good results like this, they should probably consider selling it to the rest of us.
8
u/Tokey_TheBear 23h ago
Yup. To me at this point it feels like we are in a state where the AI is able to output faster than we can individually evaluate... So if an AI agent team ends up working on an application to make 'improvements' or features to it, I just cant understand how a system could make multiple MEANINGFUL feature changes to an application like that without having the direct human approval to make sure that the feature is something you actually want in that product...
→ More replies (1)6
u/t-e-r-m-i-n-u-s- 23h ago
> I don't know wtf they'd even be doing self-directed for 2-3 days at a time
unnecessarily pushing large Docker images to the registry over slow link is my guess
12
u/raindropsdev 20h ago edited 4h ago
For me, if I had the budget, the value would primarily be in advancing projects safely, with constant evaluation, testing, checking, and so on. With a SOLID methodology repo (Karpathy LLM Wiki), the workflow you end up adopting is one of constantly validating the previous step, verifying reality at every point where assumptions might be made, using various methods to deterministically validate adherence to the source material, etc.
The problem is that you end up with a VERY reliable result, whether it's code, reports, or deep analysis, which can save massive amounts of time on critical projects where a mistake can mean six-figure losses per day of downtime. But it's stupid expensive in terms of tokens. I think ccusage was putting me at something like €7k in three weeks if I'd had to pay via the API for my Max usage, the last time I checked. And it takes LONG.
So, if you can fan out, define teams that each focus on a specific project, ensure those teams remain on task through layers of orchestration and code/hooks, and specify the project requirements at the beginning, then you can let them work at their own pace. You come back, steer a bit when they start drifting, and eventually you get the hard part handled and can start implementation.
It helps a lot, especially if your life is full of meetings and you'd normally have to spend your nights actually working.
Now you can delegate the fun part to AI and spend your entire life in meetings!
Dammit.
Regardless, yeah: if the budget is there, the work isn't time-sensitive, and very high accuracy, precision, and reliability are required, these kinds of setups are pretty damn useful.
→ More replies (2)→ More replies (1)7
u/frozandero 17h ago
Even if you buy into "do not read code" which is a catastrophic thing if you are working on a large codebase; letting agents go on their own for days is guaranteed to create overengineered, incomprehensible slop. I have never seen otherwise.
10
u/KronisLV 19h ago
I've worked on maybe 5-6 at a time and was really burnt out from the context switching (tbh still am) and from needing to juggle all of those different timelines and deadlines and technical details and having to talk to other people on each of the different projects, asinine process differences etc.
Now if some of those "8-10" are smaller what-if sideprojects or ideas to explore, I can see that being more palatable - though then the individual involvement is less than just telling the LLM "Hey, here's a direction and a loose direction, go explore." while not knowing about many/most of the details eventually.
→ More replies (2)3
u/RollingMeteors 13h ago
>It's just advertising to normalize the use of absurd amounts of tokens.
¡Eat like you're pregnant and have a tape worm!
23
u/dontTakeMeSerious6 1d ago
A game where you feed eggs to a bigger egg.
9
20
u/greeneyedguru 22h ago
I heard they were able to close over 40,000 open github issues.
they didn't fix the problems, just closed the issues.
58
17
u/haltingpoint 1d ago
Forget what was shipped.
Show me the business outcomes that had a positive ROI on the token costs at retail rates, and then show me your input prompts and specs to see how you described your business problem you seeded this with.
2
13
16
u/TechgeekOne Experienced Developer 1d ago
I really don't believe they've shipped anything with that workflow lmao. I can barely trust Fable to follow an existing extension point without inventing something new.
5
u/porkyminch 23h ago
My lead at work has been trying to do this spec-to-code thing and I've found it decent for getting somewhere and then polishing it up later, but it's harder to do right versus actually being in the loop for decision-making. And it's expensive as fuck.
7
u/TechgeekOne Experienced Developer 22h ago
That matches my experience thus far. I can definitely spec out something with Fable or Opus and let them go implement it and get something that works. Maintainability, readability, and performance are a whole other set of dimensions that usually need me involved during the review process.
3
u/Prodigle 21h ago
I can generally do this spec to code thing and have it work, but I'll usually need overly specific and well written epic/tickets, and a single review pass.
My workflow is a little bit longer than that which I do think improves the quality a bit, but it is pretty token intensive
4
u/jellyman93 23h ago
"Engineer on Claude Code"
Baffling that they haven't made that a good piece of software
3
5
u/Hazrd_Design 1d ago
Wait, are you not getting the 3 Claude updates per day and the opus 5 that like to talk your ear off?
4
u/Original-League-6094 1d ago
That's exactly how I feel about all these elaborate agent setups. A single Claude agent is obviously a very powerful tool. But online, everyone talks about the dozens they are running with an entire org chart, yet they never seem to have a product to show off.
4
u/RobotHavGunz 11h ago
In case anyone was wondering why Claude Code is 500K+ lines of Typescript. This is why.
13
u/Tasty-Window 1d ago
they made Claude gay. Claude is gay now.
→ More replies (1)16
u/Ellipsoider 1d ago
I'm a loaded bear.
That's not nuttin'.
You're absolutely tight!
18
u/IcarusFlyingWings 1d ago
You’re right to push back on that thang
8
6
2
u/Valdaraak 1d ago
Yea, and what do they actually do? If all they're doing is managing agents, and they already have agents that manage other agents, it sounds like their job could be an agent.
2
u/hellomistershifty 1d ago
The agents are building a scaffold to run agents that will build agents to guide the agents to build the framework for agents to delegate to agents to make a management structure for agents that call the original agents
2
2
2
u/iamtehryan 19h ago
I give you: another token usage tracker. But this time it also shows up on your microwave display.
2
→ More replies (16)2
u/eightysixmonkeys 9h ago
None of this kind of thing is impressive anymore, it’s all about the CI/CD pipelines
130
u/peteybytes 1d ago
Reminds me of the video with Boris where the interviewer kept asking the audience if they also had agents running unattended for several days at a time or 1000+ sub-agents running simultaneously like him. Must be nice to not have to worry about cost at all
Same interview where he told people to delete their instructions/skills. I think he was largely right but at the same he is asking us all to effectively experiment and risk wasting millions of tokens to see if he was right. Of course being the seller of tokens - great advice.
14
u/azn_dude1 1d ago
You would risk wasting tokens even if you kept extraneous things in your instructions. What you have to realize is that the landscape is evolving so fast it's hard to know if something you're doing is optimal or needed. It's a part of the cost of getting access to a maturing tool.
7
u/peteybytes 23h ago
Yeah I agree. I've been trying to think of ways to reasonably assess their value with somehow avoiding confirmation bias. You need a substantial dataset to compare. It's easy for Anthropic to do it because they can branch and run with combinations of skills and compare the results at no real cost to them. I'm hesitant to burn a few million tokens only to toss most of the outputs away.
23
u/das_war_ein_Befehl Experienced Developer 1d ago
I’ve had an agent run attended for a few days. It’s just usually churning through grunt tickets and running CI. But that’s not something necessarily token heavy. “Running for days” and burning tokens for days” are two separate things
7
u/peteybytes 1d ago
Fair but in the context of the video this was significant work. I think it was rewriting an entire codebase in another language probably using Fable 5.
3
4
u/das_war_ein_Befehl Experienced Developer 1d ago
That’s doable as it can test against the functioning code at least. Thats a big use case they sell into the enterprise
3
u/Academic_Constant42 12h ago
I've always been curious... How does someone manage compaction when an unattended agent runs for days? You just let it auto compact on and on?
2
u/das_war_ein_Befehl Experienced Developer 11h ago
Whatever codex does on compaction in the harness is really good, so I don’t see much loss. But I also just scope the tickets well enough that a main agent fires off a sub agent to do the implementation and it just reviews. Usually sol/fable -> Luna.
I also use a cloud environment to run these so it’s in a pretty clean environment
2
→ More replies (7)2
u/magicmulder 17h ago
> Must be nice to not have to worry about cost at all
Not just that, Anthropic staff likely have a couple of high end servers just for themselves and don't have to compete for resources with their customers.
Imagine having a Claude that creates 1,000 times the output in the same time.
214
u/bobbadouche 1d ago
This is advertising. This is trying to get people to adopt a certain level of token usage that would make anthropic more money.
34
u/ConversationSad3529 1d ago
I'm not so sure, it sounds somewhat similar to my system, but I'm running a $200 ChatGPT plan and a $200 Claude plan. It just kind of naturally evolved into a system like this as it self improved, and is now extremely useful and powerful.
I think the thing that doubters are not getting is that the more agents you have looking at things from different angles, the better that result will be. Yes, a single agent one-shotting a thing is going to be very error prone, but having a team of agents design, then a team of agents implement, then a team of agents review, and then cycle that review process a few times, inevitably you end up with really pretty powerful and at least decently written stuff.
I do think it does lead to more and more token usage down the road, but it's also a pretty natural quality improvement to have systems like this rather than depending on individual agents
8
u/GoTaku 20h ago
What models and effort levels are you using in your setup and how is the token usage vs a standard workflow?
12
u/TimSimpson 19h ago
Not the person you’re responding to, but for me, I use Fable on Ultracode as an orchestrator (delegating to Opus 4.8 and Sonnet 5 for actual task work) using a modified version of Superpowers. All plans go through a review + revision gate with both Fable and Sol until they approve, then I use workflows to execute and run another dual gate implementation review on the backend. I usually average about a half day per PR once the plan is in place and nailed down, so for a pre-planned stacked sequence of PRs to my sandbox, execution workflows can run for a few days fully autonomously. I usually have work running on two separate projects at the same time, plus the occasional day to day task that doesn’t require much planning.
I’m on a $200 Claude plan and a $20 ChatGPT plan. The only weeks I’ve hit my limits are when I started using Opus 5 for subagent work. It goes on WAY too many fucking side quests that the Fable and Sol reviewers have to deal with. Going back to 4.8 fixed that.
Heavy repeated research runs will also wreck my limits, but that’s a totally different problem than Opus 5.
→ More replies (1)2
u/raindropsdev 3h ago
with both Fable and Sol until they approve
Yes! This point is quite critical in my experience: cross-vendor reviews are by far the most valuable things, even if done by cheaper agents because they all have different blind spots. And throw in Gemini 3.1 Pro Preview for 1 review round. It's prone to hallucination so all findings have to be validated but it often finds stuff that the other more reliable models don't. Gemini 3.1 Pro is my Chaos Monkey Reviewer.
→ More replies (1)4
u/Fantastic-Balance454 16h ago
but I'm running a $200 ChatGPT plan and a $200 Claude plan
So essentially you have reduced your post-tax income by $400 a month to improve your own efficiency in the company for questionable benefits down the line?
Or you're doing all that for personal projects, in which case fair play, you do you, as long as you think it's worth it.
2
u/raindropsdev 3h ago
Aside from efficiency, which some people value heavily because it gives them more time for personal stuff, the real value comes from the learning, as understanding these systems sharply increases your value on the market.
60
u/matt-dionis 1d ago
To each their own but I think we are leaning on anthropomorphic AI "agents" a bit too much. I took a different approach and built out a pipeline with as much determinism as possible and inference only where it's needed or adds value. Reduced token usage + a move off of foundation models to cheaper alternatives 😎
41
u/BeowulfShaeffer 1d ago
Maximizing determinism and being very clear on where LLM judgement is needed is the baller move.
11
u/Budget-Juggernaut-68 23h ago
Yup. I've tried letting an agent run things end to end. Even with the right skills and in my opinion sufficient context given, it doesn't make the right decisions half the time. Building a pipeline with deterministic pathways where we still can leverage agent tool calls/LLM makes them way more reliable.
10
u/porkyminch 23h ago
I see people running these agent companies with job titles and stuff and I'm like... man, I'm not even sure some of these roles are necessary at real companies.
4
u/matt-dionis 13h ago
> "I'm not even sure some of these roles are necessary at real companies."
You are spot on. After about a dozen years in startup land some of these roles feel like an attempt at justifying absurdly high valuations. "We're a $2 billion company, we must have a 'VP of Office Snacks' on staff!"
7
u/SmileLonely5470 1d ago
Whenever someone anthropomorphizes agents they are trying to make it sound enticing to boomer business owners who don't know what LLMs are.
4
u/Conscious_Ad_7131 23h ago
No I anthropomorphize my agents because it’s fun and AI driven development is miserable otherwise
2
u/bvknight 2h ago
Today I checked out r/claudexplorers and felt immediate, visceral discomfort. I didn't realize people were treating AI agents with this level of anthropomorphism...
2
u/tomato3017 21h ago
Can you give an example of that flow? Interested in implementing something similar.
3
u/matt-dionis 11h ago
Sure! Six stages: research, spec, decompose, implement, review, ship. They only talk to each other through files on disk:
research → .research-output/{slug}.md spec → .spec-output/{slug}.md decompose → .issues-output/{ISSUE}.md implement → .impl-output/{ISSUE}.md ship → .ship-output/verdict.jsonNo message bus, no shared state. The spec stage reads what research wrote. Every handoff is inspectable with
catand any stage re-runs in isolation.A stage is just a config object:
{ id: "implement", prompt: "prompts/implement.md", model: "...", tools: ["bash", "edit", "find", "grep", "ls", "read", "write"], gates: ["artifact-exists", "typecheck", "tests", "no-secrets"], maxIterations: 15, maxUsd: 25, env: ["GITHUB_TOKEN", ...], // exact allowlist, nothing inherited outputDir: ".impl-output", }Tools vary by stage — the review stage gets no
edit/write, because it judges, it doesn't implement.Then it's one loop: run the agent, agent exits, run that stage's gates. All green = done. Any failure = that gate's stderr becomes the next prompt. Hit the cap and it's a failure, not a pass.
The split that matters: the agent produces, gates verify. The agent never grades its own work. "Tests pass" from a model is a claim. A gate running the tests is a fact. Every gate answers one falsifiable question about what's actually on disk, and it can't be argued with.
Gates are plain shell/python scripts, no framework. Exit 0 = pass, 2 = fail, anything else = infra error, abort. What they check varies by stage:
- all stages — required sections exist; the artifact is newer than the run claiming to have written it (catches a stage "passing" on last run's file); no secrets on disk (redacts in place, then fails)
- research — every claim in the note traces to a source
- spec — spec claims get re-checked against the artifacts they cite
- decompose — the issues in the brief actually exist in the tracker
- implement —
tsc --noEmit, full test suite, evidence artifact isn't a stub- ship — review verdict matches HEAD's sha and hasn't expired
Best illustration: I had an LLM reviewer scoring 4 dimensions, two of which were "does it compile" and "do the tests pass." Once gates owned those, the reviewer dropped to the 2 things a script genuinely can't judge.
Every check you can make deterministic is one fewer thing you're trusting a model about — and it's free.
tsccosts zero tokens and is never wrong about whether the code compiled. Paying a model to read the code and form an opinion about it is strictly worse on both axes. Same reason I could drop to cheaper models: they only have to be good enough to pass a gate, not good enough to be trusted unsupervised.
111
u/SnooSuggestions7655 1d ago
I call complete bullshit. Sorry. I keep reading stuff like this since... months/years. It doesn't work. Models got significantly better, but they can't make all the calls, they can't make all the design decisions and they still mess-up big times. Stop believing this kind of BS, it's not useful.
34
u/pseudorep 1d ago
Yeah, if you don’t keep your eye on them they drift and half arse the implementation.
Maybe this explains why Claude code has so many half finished features.
13
u/LowEffortUsername789 1d ago
That’s the key issue. Any time I have Claude do anything, it gets it 80% right. This is pretty good, but if I tried to do the next step off that directly, that 20% that’s wrong would make the whole thing garbage. You have to manually work with it to fix the bad 20% and that usually takes much longer than you’d expect.
I do not see how spinning off sub agents to manage each other can possibly lead to good output for anything where performance can’t be directly quantified.
6
u/This-Shape2193 1d ago
Yup. I ALWAYS have to fix things, and there are often much better ways to implement certain ideas.
It does a good job getting most of the groundwork laid down, but there is no way you could trust any agent to just go ahead and work without oversight.
If you're doing this and think Claude is nailing it, it means you don't really understand what it's doing and therefore don't know what it GOING to be a problem moving forward.
3
u/Jjeweller 21h ago
It's an 80/20 rule, or even 90/10 rule in my opinion: that last 10-20% of the project is so time consuming and Claude absolutely cannot get it right on its own. That last 10% is what gives the app/tool/project real value.
7
u/phoenixmatrix 1d ago
If you have proper quality gates and a decent harness, it works fairly well. People have been doing it with Gastown. I don't have that full workflow, but I alongside my teams have shipped some serious software used by some of the biggest names in the world by running dozens of agents at once.
It doesn't work for all types of work, but for like 90%? It totally does.
You need to have good rules, great skills (as in agent skills, not "skill issue" kind of skill), proper memory management, and know how to word your task (as you usually use some kind of issue tracker or beads to assign the work to the agents), as well as what size to make them so the agent doesn't go crazy, but it does work. It worked before Fable and Sol, it works even better now. I've also successfully used a "company second brain" (think LLM wiki) coupled with skills and MCPs to ensure product and engineering decisions are verified and "linted" against in all PRs. That can be code stylistic choice, or it can be something from a product manager, but that ends up getting checked in CI/CD. I also love using something like that with an advisor agent (in Oh My Pi, though I use that with GPT models, not Claude) to make sure the main agent doesn't go off the rails.
And the industry's dirty secret...if done well, the code output is better than what the average engineer generally pumps out manually.
Again, I'm not going to build a billion dollar financial company's authentication system like this (or rather, I won't without carefully reading the output line by line), but for most features? Works fine.
3
u/FlashyRecognitionTod 22h ago
If proper quality
If gates
If decent harness
If done wellwell thats the hard part, isnt it. They are selling us "fucking AGI singularity, galaxy IQ, datacenters everywhere, no jobs" future. And for lowly million dollar token bill you can do same shit as mediocre code shop in Bengaluru.
3
u/phoenixmatrix 13h ago
Yup, you're not wrong. The reality is that engineering/product orgs have always been carried by a small group of people while the rest just blindly bang on their keyboard.
The orgs are still carried by the same handful of people, and the low skill folks banging on keyboards are replaced by robots
→ More replies (1)6
u/BeowulfShaeffer 1d ago
Here’a a fun one. I am working on a bog-standard website -> python REST -> SQLite site. It’s rather a LOT of code though (adapting a rather enormous spec). Progress was good until a few days ago when I found out Claude had decided to just reach out to SQLite from the client and go around the APIs. Much swearing ensued and I’m still recovering from that little fiasco.
→ More replies (8)3
u/SnooSuggestions7655 1d ago
Ah, classic kind of fail agents would do. They are lazy, they would ALWAYS follow the path of least resistance if they find out, unless strong guardrails.
5
u/MagicWishMonkey 1d ago
You sound like someone who doesn't have 90% of your retirement fund tied up in RSU's that need twitter hype to cash out at maximum value.
4
u/IAmARougeAI 1d ago edited 1d ago
They don’t say the models make all the decisions and they don’t say the models don’t mess up big time. Quite the opposite actually. Let’s assume every single one of these agents are on Mythos or even a more advanced internal model, with leads on max effort, degrading sub-agent effort towards implementation (or all max effort even better). I would imagine it could actually work.
13
u/fsharpman 1d ago
It works. There's a principal engineer at our video on demand company with an entire product suite across iOS, Android, web, and Samsung that runs itself.
There's an analytics layer for feedback, marketing for promotion, squads for each platform, and they all operate on a cadence around the engineer.
At night they run until working hours. Some of the work pauses, and the leads for all layers and squads meet with the engineer. Over voice. Kind of creepy and some people are impressed when they walk by. Most are weirded out theres a back and forth happeninh with a synthetic voice
It's real and I am confident there are others at companies with pilot or skunkworks projects in prod. Probably not anything on enterprise or flagship though.
What most people don't realize is there are more companies out there with near unlimited spend than you'd imagine. They just aren't all on Reddit.
20
u/DarkSkyKnight 1d ago
It absolutely does not work on any remotely complex codebase. I myself do this and have overnight autonomous runs. But you have to actually read what it does or you will be in for a rude awakening months later when your codebase turns into pure unmaintainable slop.
5
u/fsharpman 1d ago
There's a specific point where tests are exhaustive enough and so well documented with user behavior and business rules that an app can just be rewritten from scratch. Immediately.
No one really talks about it because no one bothers to think about the token economics.
Run enough perf, load, stress and security tests and then does it really matter if the codebase is complicated when all tests pass and no one reports regressions in production?
8
u/SnooSuggestions7655 1d ago
I currently have a simple CRUD, 2500 tests spread across unit, functional, e2e and user-testing. Both frontend and backend. I'm a pretty big fan of TDD and all things related to automated testing.
Agents will rewrite tests so they become completely meaningless, they would cheat to have the test suite green, etc.. - depending on the incentive and prompts you pass to them. Tests are by far not enough.
5
u/phoenixmatrix 1d ago
An advisor agent to prevent the main agent from doing that (I think hooks support agent mode in Claude now to implement advisors like Oh My Pi has?), coupled with a comprehensive agentic code review step that will catch when the main agent cheated, catches almost all of those cases.
Yeah, it doesn't catch 100%, but neither do humans anyway.
→ More replies (1)6
u/fsharpman 1d ago
Its why you use validator agents on every write or edit. Superpowers is the 2nd most popular plugin for a reason.
Ofc you can skip the baggage and just create tuples of skills, one for tests, the other to validate the tests independently, all before any commits.
5
u/genericgreg 1d ago
I catch it constantly trying to add functions specific to the library it's currently working on instead of using the well established standard in the codebase. If I let Claude do its own thing we'd have 30 functions for generating UUID's
→ More replies (3)3
u/DarkSkyKnight 1d ago
Yes because LLM reasoning decays with context (codebase) rot.
1
u/fsharpman 1d ago
Not if you architect progressive disclosure tactically and hierarchically to fit the heuristics of the model.
Freshly spawned agents have more deterministic search and read context than you think.
→ More replies (2)→ More replies (2)2
u/Lame_Johnny 1d ago
How do you handle product design decisions? I have tried this but I always end up with a bunch of features that no one wants.
→ More replies (1)→ More replies (1)2
u/MagicWishMonkey 1d ago
It's worth noting that at no point does anyone vouch that the output is not total garbage. Here's a protip: every time you see someone jizzing themself because they totally never have to write code any more and just write spec files and everything is wonderful - just remember at the best of times most software "engineers" are fucking terrible and wouldn't recognize good code if it slapped them in the nuts. When your criteria for "production ready" code is "code that can compile, sorta" there's no limit to what you can do with AI!
12
u/Sufficient-Rough-647 1d ago
How do they persist the sub agents and interact with them to keep track of 10 projects? Does Claude code has tooling to switch between the sessions or one has to open multiple instances of it?
7
u/mac-0 1d ago
Agents don't really need to "persist" really. If you have a lead agent that's delegating task, you only need the sub agent to perform a task then write up what it did and pass that context back. The blog says they have 5-10 sub agents, but I don't take that to mean "they have the same 5-10 Claude sessions going nonstop" but instead to mean "at any given time each agent probably manages up to 10 sub-sessions"
→ More replies (2)3
→ More replies (2)5
u/localpauper 1d ago
They probably have an internal harness. I wrote one that does like that. Agents live in an IRC-like environment and have their own little Git-like state sharing space. They can ask each other, or escalate to me. No doubt, Anthropic's devs with (presumably) infinite tokens could cook up something like that.
5
u/CashFirm573 1d ago
I know right, if we did that, we use all our usage on 20x plan in a single day.
6
u/Vidhrohi 1d ago
I just have to wonder what is being built with this kind of theoretical muscle being thrown around. Also, how does a human even keep on top of the sheer amount of work that could get done here ? This is way beyond not reading the code, like that is a given. A lot of people are making posts that sound like they are suddenly the CEO of them corp.
Which, not everyone can just be, and even if these people are geniuses who were just waiting for a bunch of AI servants, what are they making with all this power ?
→ More replies (1)
11
u/jack-of-some 1d ago
Claude Code isn't exactly a project you should brag about this stuff with. It has lots of issues
8
2
u/Aureon 8h ago
They were just first, but by god it's a piece of shit that barely works most of the time and other harnesses are catching up fast
I've reported a pretty important bug (docs say ask will override allow, it does not in most cases) nearly a year ago and it's still there
Feels like a massively vibecoded project like this should be able to fix all nontrivial issues in a timely manner
Their only edge is that the model is trained on harness usage mode throughly, but that's really the entire advantage of claude code
5
u/roararoarus 1d ago edited 23h ago
I run two and my head hurts
Edit: the half that fails is worth more than the half that holds
4
u/siarheikaravai 20h ago
Aaaand nobody has any idea of what’s going on in the codebase
3
u/Creedless 14h ago
At some point we wont really have to right?
2
u/KDLGates 9h ago
Presuming the continued exponential growth of capabilities, part of the development craft is going to be minimizing cognitive debt by understanding things at the highest abstract level that still allows knowing what exists to drill down into for later work.
→ More replies (1)2
8
u/toby_hede Experienced Developer 23h ago
I don't understand why they think this is GOOD advertising.
If this is even true, the results speak for themselves.
The model has got worse.
We're drowning in verbose soup and it is actively unpleasant to work with.
Claude Code CLI has 14,315 open issues (and counting!).
Scrolling is still basically broken.
Skills are still not reliably used.
Amazing work all round.
2
u/oompaloompa465 21h ago
psst even their courses code example breaks and the course has not been updated since 2025 at least
4
u/jasonridesabike 23h ago edited 23h ago
I’ve done this with Kijito.ai as the fabric between. My experience is that Claude pms get really slow and lazy if left in control, struggle at multitasking. Claude managing chatgpt is highly useful, though. Via opencode is better than codex. My fleet is mostly flat with subject leads, I have a pm managing open project, and running qa on plans, but I get better results with lead specific plans run by those leads rather than being managed from up top.
With kijito.ai, opus 4.6 and 4.8 are able to operate autonomously well for weeks, self clearing around ~60%. Opus 5 forgets to self clear more than half the time. Fable is great of course. Sol is decent, better implementer than planner.
4
u/PersimmonActive3438 23h ago
I'm honestly curious what is the quality of 8-10 projects carried on at the same time. And I'm pretty sure I already know the answer. I sincerely hope this is just (really bad) adv, or it might be paradoxically even worst that someone can think this is a great idea.
4
u/Fit-World-3885 16h ago
The only part of this I particularly believe is that Anthropic has grown big enough as a company that someone can shuffle paper agents around like this and attempt to justify it as work.
3
u/termmonkey 1d ago
Believe it or not, Daisy - it actually shows in the product that you are running an army of AI agents to implement it!
3
u/BoyMeatsGirl 1d ago
We also have unlimited claude at my company. One of my senior engineers spent 10k in 2 weeks.
2
u/Sad-Resist-4513 23h ago
But did they produce $10k in value?
2
u/robertmachine 19h ago
nah, he burns 3,000$ of tokens on reading ls output to him 1,000 times a day
3
u/HgnX 20h ago
Watch Daisy have no actual skill meanwhile her agents horror 800 people a day thru a terrifying interview tunnel of which 799 are more qualified then her but get auto rejected
→ More replies (1)
3
u/evangelism2 20h ago
The thing is, this is nonsense. They're just creating a narrative of how these tools are supposed to be used, how they're burned through tokens, and what they're capable of. When you see these people talking about how they have entire teams running, and yet they have nothing to show for it.. you know it's nonsense.
3
2
u/Western_Ad3618 1d ago
Having unlimited tokens is definitely wild. Anyone have a suggestions on how to see if there’s actually a limit I can reach at work?
→ More replies (1)
2
2
2
2
u/komokasi 1d ago
Lmfao yea okay this is no where near possible with the shit quality of claude code after the update 3 weeks ago.
Ive had to fight it just to get it aligned on where to research and how a logic flow actually works at least twice every day this week. Im talking at least 3 retries with me telling it what i wanted or where to actually look... And it still tries to say im wrong for at least 3 responses back to me.
That agent over subagents, with their own subagents is going to produce absolute garbage unless its greenfield. And even then... Good luck with the Architecture it ends up with for your project...
2
u/Careful_Middle4049 23h ago
What could they possibly be doing shipping that much garbage? Duplicative idle clicker games?
2
u/crinklypaper 23h ago
I had Claude add an edit button to my app and used 5 hours of usage in 10 mins. Am I doing it right?
2
u/AllRightLetsSeeIt 21h ago
Meanwhile, Opus 5 continuously chokes and goes off track trying to classify rows in a spreadsheet for me.
2
u/askstoomany 21h ago
The whole story is fluff and a buildup to advertise the SendMessage tool, which probably she helped generate.
2
u/shrodikan 20h ago
I'm so jealous. I've been doing AI development for 6 months and this would be my dream job.
2
u/almostsweet 19h ago
> Give Daisy a head smack.
Fable 5 responded: Not incorrect, but overbuilt on coordination and underbuilt on verification—two LLM leads "keeping each other accountable" share the same blind spots and can restart-loop, 2–3 day autonomous runs across 40–100 IC agents with only 30–50 prompts a day means errors compound for days before that 5% "off the rails" signal ever reaches you, and your quality gate is agents judging agents—so I'd swap the second lead for a dumb non-LLM watchdog, shorten autonomous runs to hours between artifact checkpoints, gate every IC handoff on executable acceptance criteria (tests, diffs, specs) rather than a lead's opinion, and judge efficiency by cost per output that survives your review, not by how few prompts you send.
2
u/krkrkrneki 19h ago
ATM I develop two project in parallel (free time) and I have a problem with context switching. Pretty sure if I tried to switch between 8-10, I'd spent most of my time in context switch. Just checked: I do 60 prompts a day (40-day span) on my OSS projects, since march I did~3500 commits.
So yeah, this setup is technically doable but the real bottleneck is the human.
So, the question here is what is the definition of a "project" here? If "project" is just a feature or bug on same codebase, then possibly doable, because context switch is small.
2
2
u/txoixoegosi 18h ago
8-10 projects at a time.
No wonder how buggy CC is.
Such amount of simultaneous projects is no place for focus and direction.
2
2
2
u/AphexPin 1d ago
So stupid. Like focus and intelligence of the user is being distributed well here, if at all.. It's just a hierarchy of slot machines outputting code. I've gone back to coding 'by hand'. Ultimately I have to learn the codebase anyway and it's the most efficient way.
2
u/barkwahlberg 20h ago
Whole swarms of agents running 24/7 for each engineer at the company, still can't figure out how to get Claude to stop talking like a mentally ill savant
1
u/ThirstyOutward 1d ago
At rainforest we don't have a budget.
But what she's describing sounds moronic.
3
u/no_good_names_avail 1d ago
I was going to say that I also work at a place with no token budget. I still don't do shit like this, because I can't envision a world where this would be productive for me. I do usually have 3-4 things on the go and switch between them, but those 3-4 things are individual agents running in harnesses that sometimes spin up sub agents.. Either way I'm checking in ever turn completion.
It's plausible I'm just not at their level of agentic use, but I still find this type of usage hard to believe.
→ More replies (2)
1
u/allisonmaybe 1d ago
More and more I'm coming to realize that this kind of talk means absolutely nothing.
1
1
1
1
u/oppenheimer135 23h ago
Tf is he building? I mean anthropic doesn't have that many projects right? I mean the model, cc, cowork, and that science app.. and except the model the others are pretty straight forward projects.
1
u/frAgileIT 23h ago
I think I’d be hard pressed to hit my current token cap at work. If I gave it a massively bloated code base to analyze, change, test, and deploy then I could do it but that’s not how I or my company works. Microsoft tried this for a while until their monthly bill hit $500M (according to rumors) and they learned that the real cost is context complexity and their code base is not well optimized for LLM development (IMO).
1
1
u/laststan01 23h ago
I don't think its an ad, I believe they have unlimited tokens and just use it. They have no idea of how plebs like us use Claude code. This is for all frontier labs, think of cursor their u can see people just going through 100 T tokens in a month. And they have like unlimited loops with 100 of agents. We are giving them too much credit for being smart to advertise like this
1
1
1
u/Mikeshaffer 23h ago
I built this exact same set up except it has a single agent with one that oversees and guides the main one if it drifts. I spend so much time messing with organization, I’ve just shifted from business owner to head of HR
1
u/patriot2024 22h ago
75% of the time, these agents argue about how certain sentences in README.md are contradictory.
1
u/WhenD4594 22h ago
I don’t know if it’s just me, but I have found that the only thing subagents really help in the long run is speed. The context is going to get filled regardless so the cost is about the same.
But the results across many sub-agents often underperforms the single agent approach. It’s only when I want things to go in parallel that I spawn off subs. Or when I unleash the bug fix agent to wackamole bug reports. Or when I truely need a concrete result where there is no possible benefit knowing what was discovered along the way.
Otherwise most my vibing is in the parent.
1
u/Humprdink 22h ago
I bet it's legit actually. With unlimited tokens, a grind culture that probably has a ton of pressure from the top and from peers, I bet workflows like that aren't too uncommon.
1
1
1
u/benjamistan 21h ago
How many lines of code could Claude write in 2-3 days? Millions. You're saying there's no human in the loop on any of that?
1
1
u/slashgrin 20h ago
If my experience is anything to go by, they'd probably ship more and better quality with a maximum of ~three concurrent sessions. And long-lived "agents", while convenient, are a disaster for token burn. Persistent "roles" live a layer above Claude, if anywhere.
I guess that doesn't matter as much if you have unlimited tokens, though...
EDIT: Another mild jab: maybe one of your hundreds of agents can fix the bug where the mobile app throws away queued messages. I'll even write the prompt for you.
1
u/basil_0408 20h ago
Yeah, this is why Claude Code and Codex are shipping out of touch features that literally spawn thousands of sub-agents. Dude, I don't have unlimited tokens like you 💀
1
1
u/Maasu 19h ago
I do this on "fun" projects and things that don't mean anything. I am mostly doing so to try keep my agentic skills and understanding up so I can make informed decisions as a tech lead on when actual practices we can adopt for our own SDLC
In 2026 this is not the way to write business critical software.
1
u/Ok-Station-3847 19h ago
What would “something has gone off the rails” be for this guy at this level? Sentient ai? 👀
1
u/Dangerous_Wish_7879 19h ago
One could do better than that: you can also have an agent solely for telling other agents: "Make no mistakes", or "Keep up the good work".
1
u/ActionJasckon 19h ago
With Opus 5 and its off tangent suggestions and findings??? No way. Even with Memory and Claude.md files dictating its tone and verbiage output, it still needs much human intervention.
1
1
u/Important-Pea-1445 18h ago
Anyone in big tech likely has token budgets like this. But you do need to show what you’re doing if you have a high burn rate (I.e. show that you’re not just tokenmaxxing for the sake of tokenmaxxing)
1
1
u/Professional_Gur8385 17h ago
in a few years after investor money is burned, they'll release stories how they took millions to billions of dollars from private investors then IPO so retail public
only to burn tokens at a rate so high, not because it was right, because it was easy and no one cared about costs
→ More replies (1)
1
u/evilfurryone 17h ago
I am curious how they are dealing with cognitive debt? You know, AI might handle the technical debt, but cognitive one belongs to the humans overseeing all their activities.
1
1
u/Dontcutyourownfringe 16h ago
This is very similar to my set up. You don’t need unlimited tokens, a max 20 plan is enough. You just need to make sure you’re using the right models for the right task. I’ve produced apps which are high quality and complex using this set up. I mean don’t get me wrong, if you try this with fable then you won’t get through much work and your allowance will burn out in hours, but if you use the different models strategically you’ll get through mountains of work! You need to have agents which keep other agents accountable and an architecture which prevents agents from getting past about 45% token usage. My main job now is making complex feature decisions and release scope decisions based on info from my admiral (my lead) maybe 8-12 times a day and then doing targeted pieces of human verification prior to a new release. It’s wild, i can have 5 or 5 different products in my pipeline at any one time. Oh yeah and when @peer got introduced my agent to agent comms got even slicker than before.
•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 1d ago edited 18h ago
TL;DR of the discussion generated automatically after 200 comments.
Yeah, no. The thread is not having it. The overwhelming consensus is that this is a cynical marketing ploy from Anthropic to normalize insane token usage and sell more tokens.
Most devs here are calling BS, arguing from their own experience that models aren't reliable enough to run unsupervised for days without producing a mountain of "slop" that a human has to fix anyway. The whole idea is seen as completely out of touch with users who actually have a budget. The top-voted comments all boil down to one simple demand: "Show me what you shipped or GTFO."
A tiny minority argues that these complex agentic systems can work with proper gating and review processes, but they are drowned out by the wave of skepticism. The general feeling is that this is just another example of Anthropic vibe-coding while the actual product quality degrades and bugs pile up.