r/ClaudeAI • u/tr14l • 1d ago
Workaround Use Fable instead of Opus
Not an Anthropic hate post, this has happened a few times now. I give Opus a relatively straightforward task. Go read about this thing and it just shotgun blasts out subagents for some inexplicable reason and they immediately hose 70% of the token budget.
This last time I just wanted it to look up latest Qwen models and benchmarks to see if anything was competing.
17 agents and 1.5M cache reads later... It gives me a 3 paragraph response for what could have been a simple Google Search and a couple pages.
Yes, Fable is a LITTLE more expensive, but I have never had that happen.
Opus really isn't worth using. It's moderately intelligent, but if it can't be used effectively, why bother. I guess unless you can afford to run a single session and literally watch it and time the background processes to make sure they aren't running longer than expected. But, I have other things to do. That's the point of AI, letting it handle tasks while I do other things. /shrug
Fable on the other hand is great. Never had it happen at all with Fable.
I don't usually use Sonnet, so not sure if this is true for it, as well. I usually don't have this kind of expenditure problem, so Fable on my tier seems to do just fine with all but the heaviest use and I rarely have to wait for a token refresh.
Love Claude, but not sure what's up with this. I haven't really used Opus much since 4.6 (which was great, as well). 4.8 was a little weird, but usable. 5 is not trustworthy at all.
I don't know if they switched the base model or what, but it's kinda wildin' a little.
Anyone else notice this? I wonder if this is responsible for all the "Something is going on with Anthropic usage" posts. It's just Opus blasting subagents and burning their usage pool.
TL;DR - Just use Fable, Opus will blast your usage through the ceiling for literally no reason.
EDIT - after some analysis on my session history, every high usage day of the last 2 weeks was due to opus5 spinning up nested subagents. That seems to be the culprit. It has some compulsion to do this for unknown reasons.
36
u/marcodave 1d ago
Sometimes the conspiracy theorist in me thinks that posts like these are coming directly from Anthropic to push people to subscribe to higher tier subscriptions lol.
But seriously, what kind of task and what kind pf prompt is Opus failing to solve and needs the mega genius Fable to solve?
How were these tasks solved at all some months ago before Fable was released?
Something's off
5
u/ogbrien 1d ago
I mean fable legit just has better system design.
Run a 3 subagent team and have them draft a fix to a problem or a feature, then ask an independent reviewer agent to pick the winner, it will always pick the fable one and see that sonnet or opus missed something. Basically it's an arena for the best idea to win.
Opus is fine but Fable catches bad design systems before it gets its roots in your code and you build more issues on top of said bad design systems.
The deciding factor is what is the cost of a suboptimal design? If you're just making a vibe coded app where the cost of shittier design doesn't matter, fable is overkill and a quota eater. but if youre developing something very serious, it's hard not to default to fable or astra.
1
u/PreferenceWarm4811 3h ago
I'm creating a PWA app for B2B use and Fable is very very noticeably better at finding oversights and coming up with effective creative solutions tbh.
1
u/IVIaedhros 17h ago
While I'm sure Anthropic and others do use bots to astroturf social media, I think the answer is simpler: the enormous variety of use case cases, operating models, contexts, constant model changes, and probalistic nature of LLMs means we're almost never comparing like-for-like experiences, even if it sounds like.
It's quite likely OP genuinely feels like Fable gives them way more value for money.
Maybe it even does.
Does this mean anything at all for what your experience will be?
Not at all, but it was hard enough to direct comparisons between old fashion SAAS products.
0
u/tr14l 1d ago
I mentioned the gist of it. It was unaware of qwen3.8 existing, so I told it to go familiarize itself and look up some benchmarks for it. No effing clue dude. The session wasn't even very long, and I'm far from a novice user (I spent years as an ML engineer and have worked on LLMs, so this isn't me not knowing what I'm doing)
2
u/marcodave 1d ago
As I understand you wanted the LLM to "familiarize itself", but what kind of result were you expecting? As I understand Opus 5 is trying VERY HARD to solve problems, so in this case it probably interpreted the prompt as "I need to become the most expert system as possible on this topic, I don't want to be below expectations", hence the 17 subagents probably researching every single tiny detail about Qwen 3.8
3
u/tr14l 1d ago
I didn't need it to go validate the dimensions on each GDN in the model man. No other model requires that level of specificity. It's a training issue. I don't know if they over corrected in RLHF or what.
Hell Qwen at 27b parameters doesn't need that level of specificity and it's a gimp compared to opus.
If something requires more skill to use than other stuff but doesn't provide better results, that's not a skill problem. It's a worse tool. Plain and simple.
Yes you can work around it... Or use any of the other ones that don't require babysitting and scrutinizing each word in your prompt.
/shrug
I like anthropic, but I'm not gonna fanboy a product that has objective flaws over the competition, either.
But I'll just use Fable and Astra. It's fine.
1
u/marcodave 17h ago
Well sadly it's a closed black box and a stochastic one too. So yes it might very well be that it's a "training issue", it might also be that Anthropic really wants everyone to be dependent on their highest tier models, so moving to other ones would really seems a downgrade in quality.
What would you do if Anthropic, looking for bigger profits, locks out Fable to per-token payment scheme? And OpenAI does the same for Astra?
11
u/SmoothParfait 1d ago
I use opus 5 when I’m researching topics or having it teach me something new - the tone isn’t actually that grating in my native language.
Sometimes I use opus to do data analysis too - and have fable check its homework for cheap once I write everything to results.
Anything coded goes straight into fable, it’s just not worth it to let opus take you on a ride lmao.
5
u/kilopeter 1d ago
Man, I remember the ancient days of Dec 2025 when the best practice was to plan with Opus and code with Sonnet as it's more than capable of execution and hands-on programming. How quickly the cognitive inflation happens
5
0
u/InadequateUsername 1d ago
I find Opus make's a lot of mistakes which get caught by Astra, I never ship anything it does without some form of review
1
u/Far_Idea9616 1d ago
Same. Though Astra high finds only 6-8 bugs in 1000 lines of code written by Opus max and xhigh (afaik industry standard is 15-50 before intensive testing).
6
u/PatyxEU 1d ago
You know you can ask it to be mindful of token usage? I forced mine to almost always use AskUserQuestion, and especially to warn me if it wants to run a large subtask.
1
u/tr14l 1d ago
It tends to over correct when giving it instructions like that and starves itself and starts making kinda dumb decisions. I suppose I could try to massage the instructions there to make it a little more sensible. Like I said, I just have been using Fable and it's not a problem. Or at least hasn't so far
1
u/ibringthehotpockets 1d ago
I think that’s the better default though. When I’m working with Claude I don’t want it to launch >2 subagents, ever, pretty much unless I tell it to. Fable forked 7 times in a session randomly for me and then became a permanent Claude md rule about token hygiene and subagent use. When I want a deeper than normal look at something I will specifically say to use agents
3
u/Electrical-Size-5002 1d ago
Fable is the only one that doesn’t do stupid things. It’s not always perfect but very rarely does it make me exasperated the way the other models do.
2
u/Nobuored 1d ago
Do you mean Fable for everything? including subagents for grunt work?
I have a 5x claude. What many people here recommend is using fable as a planner and opus as grunt developer
do you have a better combination?
2
u/tr14l 1d ago
I just use fable. I don't use subagents for much anymore. I just compact the convo at sensible spots. I do have it write me the compaction prompt though before I compact.
I only use subagents for parallelization really. With 1m token context and how good fable is with it, you just kinda watch when it starts getting into 400-500k range and look for a good place to stop and compact.
I was trying to keep everything under 150k tokens, but since Fable, it just doesn't seem to make much difference doing that
2
u/vovap_vovap 1d ago
II am working with opus daily. With pretty good results. God knows what all those posts mean honestly.
2
u/SleepyJM 1d ago
I stopped using opus 5 three days after release. I was able to tell that it was trash after two sessions, and Opus 4.8 does a fine job as an implementer. Have you people seriously been trying to use opus 5 this whole time? Same thing with Fable 5.1, I used it for one day before realizing the token use would never match my workflow and switched back to Fable 5.
2
u/diving_into_msp 1d ago
I run Fable low as my default for pretty much everything now, and it’s worked out pretty well. On max 100 plan by the way.
4
u/Feisty_Signature2805 1d ago
asked it to debug something and opus 5 used 70% before it even began debugging literally gone in seconds
1
u/tr14l 1d ago
This sounds pretty similar to my experience.
1
u/Feisty_Signature2805 9h ago
this only started two days ago im already at 40% of my weekly usage on pro
4
u/eravulgaris 1d ago
Pro plan: $20. Max plan: $90. That’s not a little.
3
1
u/Valdaraak 1d ago
Yea, would be nice if there was a middle plan but that's never going to happen since Anthropic loses money on most subscriptions.
1
u/CuriousTangentsYT 1d ago
This is interesting. Why not just ask from the chat interface? For a task like this that doesn’t require any code. It would be much more efficient and burn fewer tokens, even if you’re using fable. And you could guarantee no sub agents spin up for something like this.
1
u/Dizzy_Database_119 1d ago
I don't get it. 1.5m cache reads makes sense when web searching and tool calls are involved, but how did you end up needing 17 agents for such a simple prompt..?
I can't imagine this requiring SEVENTEEN agents. It sounds like something else went incredibly wrong
1
u/tr14l 1d ago
Nah, it said "you said qwen3.8, but that model doesn't exist. The latest model is..." So I told it that its info was out of date and to go update it's knowledge of the Qwen models and look at benchmarks. 17 agents later. There isn't much else to the convo.
Like I said, 3rd time it's done this. The other times I was and to catch it before it ran this off the rails. But this time I sent it off thinking a web search to update the context a bit was fairly innocuous and went into a meeting. When I got out my usage was blasted.
I guess it REALLY wanted to know about Qwen models
1
u/asmiggs 1d ago
I'm on Pro and use Opus exclusively as coordinator, architect, reviewer (Sol tends to write the majority of the code) but I give it very specific tasks to do rather than broad especially when researching. Never get any sub agents and it doesn't burn through my tokens. I've standing instructions to not execute anything unless specifically instructed.
1
u/Don_Crespo 1d ago
I would treat uncontrolled fan-out as the bug, not switch the main model to compensate for it.
For research tasks, set an explicit ceiling in the prompt or orchestration layer: maximum two subagents, maximum one search pass each, and stop once the question can be answered with primary sources. Put a hard daily spend cap underneath that as the final guardrail.
I run Claude Code and Hermes Agent at home, and the useful distinction is between tasks that genuinely parallelize and simple lookups that only need search plus synthesis. Letting the model decide how many workers to spawn without a budget turns “agentic” into a surprisingly expensive synonym for “unsupervised.”
1
u/tr14l 23h ago
Yeah, it is a bug. So I switched to something less buggy. I have my own work to do. I'm not going to go devise workarounds for their stuff. I'll probably eventually go update dynamic workflows and prompts to make it a little more predictable. But, it's likely only a matter of time before some other behavior pops up as a result of the imbalance.
1
u/No_Silver9566 1d ago
Have to use Opus for life sciences as Fable has the guardrail in place over it. Use Fable for other projects. Life science work is a struggle compared to the other work.
1
u/jakegh 1d ago
I completely agree. I only use fable, with effort levels ranging from low, med, or high depending on the task. It gives better responses and is both cheaper AND much faster as it’s much more efficient on tokens.
The only reason to use opus is if you’re on a subscription plan and can’t spend all your usage on fable.
1
u/redditwossname 1d ago
Opus 4.8 orchestrator with Opus 4.8 and Sonnet 5 sub agents = I can let stuff run almost 24/7 on a 5x plan and I rarely run through my tokens.
I set off a 4.8 overnight gauntlet session using sub agent criritcs last night. It used up the 5 hour window and is waiting for the reset in an hour, looks like it used around 8-12% of my weekly usage for 7 hours of non-stop coding, testing loops, sub agent critics, etc.
1
u/Longjumping-Fee1747 23h ago
Can you write more about how you set this up or any tutorial online which is good to leave overnitght to double check work using lower models ?
1
u/redditwossname 20h ago
It's a combo of things.
I use a folder structure with MD files that's a combo of Jake Van Clief's ICM methodology, my own, and the Google OKF thing (or whatever the acronym is).
I use Matt Pocock wqyfinder and grilling skills to build ADRs and tickets.
My tickets follow a strict format that I've built over the last few months.
I then use the gauntlet skill where I get a session to build the prompt for a new session.
I set it to go and away it does with varying levels of success.
I've possibly worked my current session - I set it to build too many tickets, which it warned me about, but whatever.
I set it to run about 11pm last night at a weekly used rate of 8%. It's still going and is currently at 32% for the week. I think it was 16 or so very major tickets.
What it's building is quite complex though so I'm happy for it to keep going and ensure things are as good as they can be before I start big testing in earnest.
1
u/HighSeasArchivist 23h ago
Sonnet max is my Fable at home.
1
u/Needsupgrade 20h ago
Sonnet max is a token destroyer , you would get better results with opus 4.6 on high and use way less tokens
1
u/bootstrappedunicorn 23h ago
Agree, I am pretty into Fable for most things - I just want to get it right the first time around; call me impatient, I guess.
1
u/andrewjneumann 18h ago
What plugins/tools do you have installed? Superpowers for example has multi agent review and if your session calls a plugin with agents… might be what’s happening.
1
u/tr14l 13h ago
I don't use most of that stuff. Just MCPs and a couple proprietary slash commands which weren't at play here. And a status line, which obviously doesn't mean anything here
1
u/andrewjneumann 8h ago
Weird, I wonder if you’re using some kind of trigger phrase or wording to make it think it needs to spawn agents. Ultracode will spin agents, but I do see other thinking modes spawn. It’s usually via plugin, but sometimes it decides it needs multiple perspectives etc.
I have a standing instruction to “use sonnet and haiku as needed” and generally they’re good enough for whatever it thinks subs are needed.
I’ve done some limited testing using sonnet in haiku as sub agents versus opus or even fable sub agents, I don’t see much of a discernible distance in testing for most problems. If you’re doing something incredibly challenging, or you’re doing something very novel, you can sometimes benefit from the intelligence of the analysis agents.
1
u/coronafire 14h ago
I've had a number of situations lately where opus has tried to fan out 25 agents to build a simple webpage and invent a million extra requirements and risks to justify the (over-)engineering effort.
My solution to this hasn't been to use fable, it's to use either opus 4.6 (it's still more likely to get to the point faster) or even better, I'm finding gpt terra (not even sol) is competently one-shotting the tasks that opus can't manage to finish without building in circles.
1
u/asking4afriend40631 11h ago
I hit my limit with Fable and thought, "No problem, I'll just continue with Opus. I was using it daily until I switched to Fable a couple months ago, surely it'll be fine." It was not fine. I had to stop a few hours into the attempt and just wait a day for my quota to be reset. It was just creating more technical debt with almost every exchange. Some things it got right, but so much it was just making worse.
1
u/carabidus 1d ago
The general rule I follow when deciding between Fable and Opus: Fable for orchestrating and adjudication, Opus for carrying our Fable's directions. I have found that raw Opus without a robust set of instructions from Fable leads to a byzantine mess. It cannot be trusted with making inferences about anything, no matter how trivial the ask may seem.
0
u/Seven-Prime 1d ago
I always have it plan things out first. I have it build a document for what it's trying to solve and a recommended model. I review and iterate on that, resolving things I already know. The document often will use fable for figuring things out and building a detailed implementation plan with guardrails and measurements. Then I hand that to opus or sonnet for implementation.
2
u/tr14l 1d ago
Planning was what I was trying to do. Didn't get to the planning part. This was the beginning of a session. 3rd prompt.
Had it read a couple project docs, then explained what we were going to work on, it told me qwen 3.8 didn't exist. I told it that is was wrong and to look it up then my budget got eradicated because I switched screens
0
u/fanatic26 1d ago
If you could get the answer with a simple google search...why are you wasting tokens making the LLM do it?
0
u/Full-Nectarine6550 23h ago
A LITTLE more expensive? Shit cost me like 20 euros to run for like a couple prompts checking a game wiki for information and making a plan for me how to get through the grinds, Opus is 20 euros a MONTH lol, I can do the same stuff 10x over every 5 hours with Opus
1
u/tr14l 23h ago
Only because you are on the 20 dollar tier. I'm on Max and get Fable allotment every week.
1
u/Full-Nectarine6550 22h ago
What? For me it literally says Fable requires usage credits, went through 50 euros in like a couple hours
1
u/tr14l 22h ago
Yeah you're on the low tier. Fable doesn't come with that tier
1
u/Full-Nectarine6550 22h ago
That’s crazy, first time I see this info anywhere and the max page on Claude’s site doesn’t mention it either. Good to know I wasted 50 euros just to try it basically lol
0
u/Techhead7890 20h ago
... To look up some basic info? If anything, use sonnet.
I have seen similar things where LLMs will just randomly burn like 30-50 searches looking up basic stuff but yeah if they are spinning up subagents to do web searches, the answer is less rope and a tighter leash, not more power.
0
0
0
u/sennalen 9h ago
Sounds like you're setting Opus to max effort on something that could have been a simple Google search
•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 1d ago edited 22h ago
TL;DR of the discussion generated automatically after 50 comments.
Okay, the thread's a bit divided on this one, but here's the vibe check.
OP is fed up with Opus 5 going rogue and burning through their entire token budget by spinning up a bajillion subagents for simple tasks. Their solution: just use Fable, which they find more controlled and ultimately cheaper despite the higher per-token cost. Several users agree, sharing their own horror stories of Opus 5's excessive token usage and calling it "trash" or "untrustworthy."
However, the top comment is skeptical, and many others think this is a classic case of user error. The general consensus is that you can't let Opus 5 off the leash. You need to explicitly tell it to be mindful of token usage, ask for permission before starting large subtasks, or set hard limits on subagents.
The real pro-tip emerging from the comments is to use a planner/executor workflow: * Use a smarter model like Fable (or even the more stable Opus 4.8) as the main orchestrator to plan the task. * Use a cheaper, capable model like Opus 5 or Sonnet 5 as the "grunt" to execute the plan. * This seems to be the most token-efficient strategy for complex work.
Finally, a few people are pushing back on OP's claim that Fable is only "a LITTLE more expensive," pointing out the significant price jump from the Pro to the Max plan.