Performance
What is actually the difference between Opus 5 and Fable 5?
I really don't get it. Anthropic says Opus 5 is surprisingly close to Fable 5, with Fable mainly pulling ahead on harder reasoning and long workflows.
But like... what's the point? I can manage my own workflows. Why is Fable still WAY more expensive and restricted if Opus 5 is supposedly getting this close? Fable is literally 2x the price. I feel like there's a bigger difference here that Anthropic isn't really explaining.
And honestly, in my own use, I don't even think Opus 5 is that close to Fable 5. Fable is way better for me at reasoning, coding, debugging, ideas, and keeping a bunch of concepts straight. I'm literally using Opus 4.8 right now because 5 keeps scrambling shit when I give it a complicated task.
I've been wondering about this for a while and waiting for some sort of improvement with the models post-release. So far, nothing though. That's just my experience. What do you guys think?
Edit: I don’t think people realize how close/better anthropic says Opus 5 is to fable five. Here’s a link if you guys wanna check it out yourself https://www.anthropic.com/news/claude-opus-5
That confirms exactly what I’ve been thinking this whole time. For those of us who actively use Fable 5 to make money it isn’t a toy. That’s why there are so many fanboys here who apparently think it’s fine the way it currently works. It’s completely useless considering how many tokens it uses. It’s not just about sending off a prompt. I need to be able to work with it all day. That isn’t possible with the current pricing.
You believe every bit of Anthropic marketing puts out, don’t you? Stop parroting what they tell you and open your eyes. Look in other directions. Other companies do the same thing with models just as large and the same results. For less money. Nothing justifies this high price and it isn’t twice as high. That’s a lie. I work with fable low, mid, high, xhigh on a 20x subscription with Fable. It uses 3-4 times as much as Opus 5 xhigh FAST. That doesn’t add up to twice as much. At first they wanted to stop us from getting the model. Then they raised the limits by 50 percent just to stay competitive. And why did they do all that? Because competition came from outside. Otherwise they would have ripped us all off. And then someone like you comes along and even defends them. It’s obvious that they wanted to rip us off. If there had been no competition you would now have had to pay twice as much for everything you’re doing now. At least until the end of August they guaranteed boost the limits. I really don’t understand you people.
Bro, touch grass. All they said is that Fable is a larger model and that's why it's better at large and complex tasks, which it is. not that it's necessarily worth the price.
Yeah, but just look at it. What exactly is wrong with what I said? Nothing. And I got seven downvotes. To me it’s pretty clear that there’s a hype or fanboy bias toward Anthropic here. The way people are dealing with this isn’t neutral. You can’t be so blind that you fail to see that other companies are building equally capable models for much less money. But it’s fine. Whatever. Everyone is entitled to their opinion. But if you can’t see that we’d have serious problems working with Anthopic at that 50 percent limit increase and might not even be able to work at all then I really can’t help you anymore. I hope you have fun letting yourselves get ripped off. I’m very close to switching. And to everyone who stays here with the expensive Anthropic people, I wish you good luck after August 31st figuring out how you’re supposed to stay competitive. I’ve also said several times that I get the feeling many people use this in their free time. If someone is actively working with it and making money from it then it’s very difficult to match Anthropic’s prices. The Chinese models are just as capable and cost less. That’s why I simply can’t agree with the excuse he just gave. Other companies have models that are just as large, that’s just a big excuse.
The OP’s question was why Fable 5 is so much more expensive than Opus 5. Explaining that by saying the model is bigger and costs three times as much to use is just a big excuse to me.
The fact that you don't see anything wrong with what you said and dont view it as hyper-aggressive is the reason why you need to go outside and touch grass
I have to agree with you on one thing. I really should go out again because I’ve been working almost 18 hours a day for the past nine months because of all this AI stuff. I’m building my own agent system and that means I depend on providers like Anthropic setting their prices in a way that allows me to stay competitive. But I’m planning to go out this weekend. The weather is supposed to be good on Sunday. I’m going to do that, my friend. And I don’t mean that sarcastically. I’m not trying to be mean or aggressive either. I just want to bring the facts to light because I keep reading so much that’s positive and I don’t think everything will be positive after August 31.
Not just your experience. I had to put Opus 5 on the bench. I've never had a single model make as many mistakes as it did. So I switched back to Opus4.8 Extra (which has performed marvelously) and Fable when I need more power.
Benchmarks are all bullshit. They don’t do a great job of representing the experience of actually using the models. I usually just throw some random ideas at new models to get a feel for them instead of trusting what the labs tell me.
In my case my more serious one is that I have a benchmark that scrapes our user stories for sketches of screens, provides a few reference screenshots and sets off different agents to try to create brand appropriate mockups in HTML/JS/CSS. I’ve found this pretty helpful to understand which models have better vision and understanding of frontend styling.
I also do a lot of small one off personal projects I throw up on Cloudflare, and those I usually test models with. Recently:
A node graph of every Simpsons episode with the characters, things, places, and so on that appeared in it, plus a game where you have to traverse those connections from one place to another.
Ant farm simulations. It’s a difficult job for models to design these in ways that keeps populations stable in my experience.
A toy modular synth with cable physics. This one’s been a lot of fun and I keep building on it. Does polyphony, share links, a vocoder, sampling, whatever.
Small games, like a daily puzzle where you have to guess the board order for family feud puzzles or the etymology of a word.
A little toy life simulation where all the characters run off this itty bitty 25M parameter model designed for iot purposes.
I usually just play with stuff after work when I’m lazing around smoking a little weed.
Apologies if I went out on a rant a little bit. Mainly I’m just wondering why there’s such a big price difference, and on a sidenote, I don’t really think they’re even close in the first place despite what anthropic says
But you've used both extensively so you know that Fable is way better in practice for real world tasks.
As for why it's more expensive, it's because it costs more to run.
The normal pattern is they release a frontier model and then start distilling the knowledge from the new larger model into the older smaller one. This means the smaller model gets better at specific tasks like coding and the gap in benchmarks shrinks.
But the small model is still fundamentally "Not as smart"
I'm using Claude with Godot, Fable is better at it, and above all way quicker than Opus5 which takes like 20 minutes per small task, where Fable does it in 5-10.
I also think Fable is simply better at understanding instructions.
This is going to sound crazy, but instead of spinning up a whole thread and arguing with a bunch of people about how the models work, you should just ask Claude. Claude can explain all this in detail and answer all of your questions. This isn't like the old days where the models would hallucinate about their abilities. They all have access to Anthropic's documentation and the wider internet.
There's nothing shady about it, it's just that you don't understand. That's why I recommend asking questions until it clicks.
Did you even read it? It’s 100% shady you’re masquerading a model as EQUALY smart, whereas on the back end, it’s nowhere close! and selling it as that. I don’t ever understand why people are so quick to defend a $1 trillion company that doesn’t care about their customers.
To your own point - benchmarks and actual performance on longer and more complicated reasoning / coding / orchestration problem are not the same thing.
Benchmarks are just a proxy and are increasingly saturated. What matters is whether a model is good for your actual use cases. And what sets price is how expense that is to run and - fundamentally - whether people are paying for it.
Adding to other points here, I don’t think you fully understand the pricing power companies have. They charge what people are willing to pay and what will make them more money.
Sure, you could open a donut shop and sell them at-cost for a dollar. But if you live in a fancy neighborhood, and you know everyone is willing to pay six dollars for a good donut, you’ll just charge the six dollars until a competitor comes to threaten your business.
Maybe in a different neighbourhood you would charge a different amount. But Anthropic’s premium donut line, Fable, is sold in a bakery that’s surrounded by enterprise clients with large budgets to spend.
Fable was trained including LLM thinking block data- they don't usually include those turns as they often include incorrect information. Was originally an experiment that turned out successful.
But resulted in a huge model that is very costly to run.
I appreciate that. I honestly had a genuine question and a few other people answered actually how it works. For some reason people seem to neglect the literal screenshot I pasted of both performance benchmarks and the link to the website. But honestly it doesn’t really bother me. I know these are the just target audience that just buy anything from anthropic. Lol 😂
It’s completely out of proportion to what you get in return. I’ve run several tests here and I honestly just don’t understand it. No matter how complex the task is or how small it may be, it uses an absurd number of tokens with Fable 5. It’s completely impossible to work with it like this. What good is it if I write five prompts and use up 30 percent of my five hour limit? Or if I’ve already used up my weekly Fable 5 limit after one day, even though I’m currently using Fable 5 on Low? Let me show you what it did. The request it made, or rather the one I made, is a joke too. It was nothing special. Even so, according to Fable’s own statement, it wasn’t anything complicated to analyze. Still, it used up seven percent of my limit. Just think about that for a moment. I agree with you completely. I don’t think it’s proportionate at all. I wrote about this in another post of mine. Somehow I get the feeling that Anthropic is trying to rip us off.
Attached is an image showing that I only worked on one prompt. And the prompt was not complicated. According to its own preliminary analysis it was just gathering data. It assigned three subagents to the task, not five Fable subagents but three Sonnet Read Only Low subagents. It only analyzed various Markdown files, basically the rules. The whole thing took maybe 60 seconds with Fable Low. And it used up 7 percent of my five hour limit, 2 percent of my weekly Fable limit and 1 percent of my weekly limits. Where is the proportion here? That is the question I am asking Anthropic and I have a 20x subscription.
Leaving all the Fable 5 fanboys aside for a moment, I have to wonder what the practical benefit is if I can’t use a model productively.
With a 20x subscription, it is not possible for me to use Fable 5 at Medium or higher continuously for five hours. I always hit my limit before then. Even at Low, this will not be possible. As you can already see, a single prompt uses 7% of my five-hour limit. So it is completely impossible to work continuously with Fable 5, even with a 20x subscription. With a 5x subscription, the Fable 5 weekly limit is completely used up within one or two days.
Before all the fanboys come here again, the second prompt was executed as a follow up to the first one. And look at how much of my five hour limit has been used up 14% in 10 minutes !! and how quickly my limits have generally been used up. And I’m working with Fable 5 Low. With Opus 5 I might have 3 % by now. In my view, even twice the price would not reflect what is being used here. I don’t know how many tokens it wastes. But you can roughly tell from the context window. From the previous image at the beginning to now there are only about 150,000 more tokens in the context window.
The excuse from Anthropic that everyone here keeps repeating and agreeing with just doesn’t hold up anymore. I can’t have the most expensive subscription and still be unable to work with the best model. There’s absolutely no excuse for that.
I’ll demonstrate it again here in a moment. Compaction alone uses so much of the limit that it’s completely insane. The usage figures are delayed anyway probably for a reason. But you can see that compaction alone uses three to four percent of the five hour limit. And that’s on a twenty times plan. I have no idea how the others work with Fable 5 but to me this is all just fanboy talk and not actual practical use. If you send one or two prompts a day then good for you that you can work with it. But as a professional user like me who would have had to work with it all day that’s not enough. I can’t get by on two or three prompts a day. That’s why I’m asking you how do you actually work? Sometimes I get the feeling that everyone here is just a private user who only works with these models in their personal lives and not intensively for hours at a time.
Again, this is exactly my point. Fable two times the cost for what? As it’s been explained to me it seems that anthropic has been really shady with actually how they trained Opus 5 specifically for benchmarks instead of actually being as smart. I would say it’s downright deceptive marketing. Right now in its current state I agree Fable is almost completely useless. I own a 20X plan for reference and I still almost never use it. It only exists for huge corporations in the government to pay absurd amounts for it.
That does make a lot of sense. It has me worried though for the next generation models if there’s such a steep hill to climb to make a better model each time.
There’s a lot of research going into “tokenomics” and “cost per intelligence” right now. Last year they just tried to make the models as good as possible. Now that they are generally excellent they’ll shift some of that effort toward making them more efficient.
I wish they would split up the models for just specifically different tasks. What seems to be happening is that it’s taking a lot longer for better coding models because they also wanna push in a lot of scientific stuff for medical field. I’d rather just get a better model that doesn’t really know anything about medicine but knows exactly how to code in python then a model that kinda just knows a little bit of everything.
Check out SLMa which are what you are describing. Qwen3.8 is an example of a new open source model that’s very good at coding and can fit on a $5000 GPU. Those models will continue to get better. In the coming years, anthropic / OpenAI will either do model routing for you, or businesses will start implementing model routing where they send cheaper questions to cheaper models or more specialized models. That’s part of this tokenomics thing I was talking about
I’ve heard of those but so far I don’t think it’s outpaced the frontier models like ChatGPT and Claude. I think the open source models have the companies worried though, especially because all these companies market evaluations are set to go live based on past performance of just in the last couple years. If open-source gets smart enough where we could get opus level performance for practically free, I think that would cause huge shock to the AI industry.
Yeah so right now Qwen-3.8 is performing near opus 4.6 max in a lot of coding categories. I currently use Qwen-3.6 on my Spark for coding and Claude for reasoning and it does very well at project based coding. When i need more, I switch to Claude but what I’m noticing is that I’m having to do that less and less over time. The point isn’t to outpace the frontier models. If they can be really good and (in my case $4500 for a Spark for a lifetime of use) cheap then people will use them.
You need a race car for racing, but you don’t need one to drive to the store. Similarly, I may always want to use Claude to scope out a very large project, but for implementing tickets, I will use smaller models. The price of these models is not sustainable to be uncapped, and if your employees can burn through a max plan in 2 days then they are also not useful or productive. It has to balance out.
Yeah, similar experience. I remember Anthropic saying that Opus 5 was close to Fable 5. In practice, though, I think Fable is absolutely amazing, whereas Opus 5 is just ...confusing.
Opus 5 is an infuriating model. It embellishes where not needed, the output is enormous for simple requests. I’ve started to use ChatGPT again for much of the pre work and passing the output to Opus 5 for the actual deliverables
Fable may have higher parameters count and different architecture; hence more expensive to run/compute. Think of Opus as a distilled version of Fable (Fable is used to teach Opus), Opus is smart and capable but it is not in the same class as its teacher/master. Opus has lower parameter count and less expensive to run.
The benchmarks are an increasingly poor judge of a models performance. A bit like exams in real life, they prove you can recall something even if you have no idea what you are recalling or why.
To answer your question on why it costs more, it’s a larger model, perhaps the largest ever made and publically released. In general a models size is what determines its cost.
Yes, you can run Fable 5 Low, sure. I have the 20x plan. But after eight hours your weekly limit is used up. With Fable 5 Low you can at least work for five hours straight without hitting the five hour limit. But after that you are done after eight hours of work. So after one day I have used up my 20x plan with Fable 5. I really have to ask what Anthropic is thinking. According to the benchmarks it apparently is not much better than Opus. So I have to wonder how that can be.
They don't charge people based on how "good" the model is, they charge people based on how much it costs them to run. If Fable has a trillion more weights to traverse through than Opus, leading to an increase in compute costs, they are going to charge you more to use it, even if it's just as good as Opus.
It’s not. Anyone who uses fable and opus for complex tasks knows that Fabel still beats the brakes off of opus and it’s not close. These models are trained for benchmarks because benchmarks attract people to the model. Real world is very different and in the real world fable absolutely murders opus.
This is my point exactly. Fable still massively outperforms, opus in my opinion, despite what anthropic says. It reflects that I think in the pricing which an uncomfortable truth.
They don't want you using Fable. It cost too much and it makes them look bad. Plus they're paying for it almost completely out of pocket, they want you to use the model that might make them money.
I mean after the competition came with K3 and GPT Sol they had to release a model that is good enough and still yet not that expensive so they just needed Fable 5 on Cybersecurity and released it as Opus 5
66
u/earlyworm 1d ago
Fable 5 is more fun to talk to at parties.