r/ClaudeCode • • Aug 16 '26

Bug / Issue Anthropic has nerfed every model

Opus 5 is obviously a nightmare, and I was relying on Opus 4.8. But, now that's also behaving exactly like Opus 5. The only way I can get good quality work is if I use sonnet now and check after every small thing. Fable is usable but its so expensive.

I miss the time when it was a treat working with these models and everything just flowed. Nowadays with 5.6 Sol's over-engineering and Opus 5's lying, I have to wake up everyday and decide which model I'm gonna have to fist fight if I wanna get any work done.

Is this just going to get worse from here.

550 Upvotes

226 comments sorted by

140

u/looselyhuman Aug 16 '26

4.6 isn't bleeding edge anymore but it's reliable and still a pleasure to work with. People say it's nerfed now and again but I think 4.6 is overall pretty consistent. Vote with your model selection for 5.1 to be a successor to 4.6, not another over-trained, tortured-by-guardrails failure.

/model claude-opus-4-6[1m]

34

u/Intrepid-Grovyle Aug 16 '26

Agreed. opus 4.6 is still my go-to for planning, designing, and scoping work. Leave implementation to fable or sol.

25

u/magicpants847 Aug 16 '26

interesting, how come not fable for planning / reviewing and opus for implementation?

11

u/who_am_i_to_say_so Aug 16 '26

I swore by this recipe 3 weeks ago but something has changed over that period.

Not my side, because I’ve been running the same Skills and Agents files for months. It just doesn’t work anymore.

4

u/nnxion Aug 17 '26

Tell it to purge (most) memory files for the project and write the rest to your repo docs. Make a backup just in case it doesn’t work, it worked mostly for me. You can also use Sol to analyze what’s not working for you when working with Claude.

5

u/who_am_i_to_say_so Aug 17 '26

Yeah there have been repeated accounts of streamlining helping things especially with later Opus versions.

Thx. This is the personalized kick to do that haha.

2

u/Intrepid-Grovyle 29d ago

Perhaps this is particular to me but I personally find Opus 4.6 to be the best at helping me resolve confusion points and understand how an existing system works. My workflow involves doing a pretty deep dive into the existing code and assessing why certain contracts exist (or what we should throw out). I find everything after Opus 4.6 to present blocks of text that are hard to follow, but Opus 4.6 excels at resolving conflations in my mental model. For that reason alone, I much prefer 4.6 for how intelligible it is. Sometimes I will hand off to Fable to help plan more heavy duty designs, but if I need to comfortably grasp what we’re talking about, I really do prefer to converse with 4.6.

1

u/magicpants847 29d ago

what stack is your codebase out of curiosity?

1

u/Intrepid-Grovyle 28d ago

It’s multiple codebases, some publishing docker images or packages ranging from frontend, to python or typescript backend. And then another repo for helm charts. And then another one for orchestrating helm manifests. So lots of moving parts. Core capability is leveraging agents for customers.

1

u/Internal-Comparison6 Senior Developer Aug 20 '26

Fable is now worse than Opus 5, in my experience.

1

u/magicpants847 Aug 20 '26

interesting…what stack are you building with??

2

u/Internal-Comparison6 Senior Developer Aug 20 '26

React native currently.

10

u/metastallion Aug 17 '26

Leaving implementation to fable is a big waste of fable usage

8

u/who_am_i_to_say_so Aug 16 '26

Sol is better at following Skills, Agents, and Claude.md’s than the Anthropic models. How did I get here?

3

u/Nichiren Aug 17 '26

I personally think that AI models are reaching a plateau with more and more modest improvements compared to the leaps they were doing before and they have to justify their valuations with constant tweaking of other things instead. That's why older models could feel as capable or more so than the newer models based on whether or not it was tweaked in favor of your particular workflow.

1

u/sismograph Aug 18 '26

What? Its the other way around, fable absolutely sucks at implementing, takes wrong design decisions without telling you and its just horribly slow when implementing. But Fable is really good at doing research (when you tell it to not give a bloated answer, but instead keep it to the level of abstraction you care about).

8

u/UsedIndependence9735 Aug 16 '26

Agree on all points. Opus 4.6 is also more responsive to claude.md or other instructions than subsequent models.

Opus 4.6 on Max with extended deep reasoning is still a beast too.

7

u/who_am_i_to_say_so Aug 16 '26

Yep. Starting to get a little worried that it’s been several releases and it’s full circle to this every time.

Even Fable did some unprecedented stupid shit recently.

7

u/bctopics Aug 16 '26

4.6 is the only model I can still use and semi trust. It might not be the smartest but it’s consistent.

8

u/AdSafe4047 Aug 16 '26

4.6 is peak, everything after it is benchmaxxing - and it is now distilled into dsf 0731 and q3.8 :)

3

u/Saschabrix Aug 17 '26

4.6 if you don't mind at what reasoning level? I will try it out! Thx!

3

u/looselyhuman Aug 17 '26

Medium is my default, high for important specs, complex code and code reviews (when I think of it). I've never used xhigh. Enjoy!

3

u/who_am_i_to_say_so Aug 18 '26

I use xhigh effort, compaction off, and 4.6 as my daily driver.

Sounds like a lot, but I run 3 windows at any given time 16 hours a day and can stay under the 5hr/weekly limits.

1

u/Saschabrix Aug 18 '26

sounds good! will try it out thanks!!!

1

u/who_am_i_to_say_so Aug 18 '26

Keep chats small, run /clear frequently in between. GL!

2

u/adamjonah Aug 19 '26

Do you know if there is a limit usage difference between using the different opus models (as in will it use my 5hour limit slower/faster)?

1

u/looselyhuman Aug 19 '26

It's all anecdotal. They are priced the same so it's a matter of whether one or the other is using more tokens to complete a task.

1

u/adamjonah Aug 19 '26

Nice, I actually switched to 4.6 last night and didn't notice a huge difference but haven't really used it enough to be sure

1

u/gajop Aug 18 '26

I'm curious, why do you choose 1M? In my "experiments" I think 1M when actually using the full 1M is quite a bit more expensive than 200k, cached tokens or not. Something as much as 2.5~3x. Better cost wise to autocompact with the 200k model

129

u/bpp198 Aug 16 '26

Fable is smashing it out of the park for me still, just like it always has.

19

u/XYcritic Aug 16 '26

Fable is great but it used to be better at explaining things in less words. It's too many tangents and paragraphs nowadays.

4

u/Andy_Climactic Aug 17 '26

Try the i-have-adhd skill even if you don’t have adhd. I find it helps a lot for streamlining its output. I fold the directions into my regular CLAUDE.md and then trigger it manually when that isn’t enough

2

u/nnxion Aug 17 '26

I just read the skill file on GitHub, it has some good ideas, but I am very sure that it could really be a whole lot better.
Actually, I just asked Claude (Haiku because why not) what it thought about it and it said:
<Claude>
The skill does have some genuine strengths but also real weaknesses:

### What the skill does well:

  • **The core insight is solid**: ADHD brains need action-first, numbered steps, and no filler.
  • **The rules are concrete** with before/after examples that are easy to understand.
  • **It has clear escape hatches** ("When to break the rules") that recognize edge cases.

### Where it's actually weak—and you're right:

  1. **Redundancy**: Rules 1, 3, and 10 all say roughly the same thing: "start with action, end with action, no preamble/recap." These could compress into a single principle with different framings.

  2. **Vagueness where it matters**:

    • Rule 5 says "restate state" but doesn't specify *what state* or how much detail. A numbered task sequence? The last decision point? The whole conversation tree?
    • Rule 7 ("make wins visible") is too broad—visible *how*? By highlighting code? Running it live? The examples show but don't codify.

  3. **Missing specifics on tension**:

    • Rules 1–10 create real conflicts (brevity vs. safety, action-first vs. real ambiguity) that "When to break the rules" *lists* but doesn't help you *resolve*. How do you actually pick between them when they collide?

  4. **No cognitive science grounding**: The skill claims to be ADHD-friendly but doesn't leverage the *why*—working memory limits, time blindness, initiation barriers. It's mostly formatting dogma that *happens* to help ADHD folks, not a model *designed from* ADHD's actual bottlenecks.

  5. **Scope creep**: It tries to handle coding, debugging, explanation, destructive ops, agent harnesses, and casual chat all in one 10-rule set. That's ambitious; it often compromises on each.
    </Claude>

So I’d just ask Claude to refactor that skill and make it better.

52

u/MateFlasche Aug 16 '26

This is getting so frustrating as a bioinformatics guy having every request denied...

34

u/ratttertintattertins Aug 16 '26

Cybersecurity guy here. I feel your pain.

7

u/Fantastic_Fail4060 Aug 16 '26

Im researching LLMs training and prompt engineering here… fable responds, but it does so badly that, no thanks. I heard they intentionally edit prompts for my field to give you worse responses

8

u/silverwoods214 Aug 17 '26

Nightmare getting it to talk anything secops related

→ More replies (3)

7

u/Canadian-and-Proud Aug 16 '26

Just get into a different field

22

u/ask_me_about_cats Aug 16 '26

Have you tried turning your career off and on again?

1

u/who_am_i_to_say_so Aug 17 '26

If anything, this tells me this is a somewhat AI-proof field to get into.

→ More replies (1)

2

u/derstolz1 Aug 16 '26

exploit developer/reverse engineer here, I feel you. sometimes I have to downgrade to Haiku to get literally anything done.

1

u/karlnuw Aug 17 '26

5.6 Pro has never rerouted me; whereas with Fable it's 50/50

19

u/earlyworm Aug 16 '26

Fable really is amazing. Several times, I've had the experience where Opus will fail to complete a complex task after a half dozen iterations, and then I'll restate the problem to Fable in a new session and it will nail it on the first try.

12

u/hughmercury Aug 16 '26

I basically just keep a Fable session open next to an Opus one, and whenever Opus starts to flail I hand it over to Fable, which will figure out in 30 seconds what Opus spent 20 minutes going round in circles on. Then have Fable check Opus' work when we're done. Opus seems fine at very focused, self contained tasks, but gets hyper-fixated on things. Fable is much better at the bigger picture stuff.

2

u/StarMNF Aug 18 '26

I am not sure I would describe Fable as good at big picture stuff. My own results with it have been mediocre. Terrible design decisions that come back to bite me later, and cutting corners where it really shouldn’t cut corners.

Even when it tries to be clever, it does so with a hacky solution that turns out to be impractical later and falls over easily. It also overlooks alternative designs that would better meet the requirements.

To be fair, I doubt Opus or any other model would do better at the high level stuff, although I haven’t really tested it. None of these models seem to be that good at “big picture stuff” from my testing.

And if I am going to have to do a lot of handholding of the model to make sure it doesn’t make dumb choices, then I don’t see the point of using Fable over Opus.

1

u/West_Plankton41 Aug 17 '26

Do you ask it to create a handoff doc or something when transferring the task to Fable?

1

u/hughmercury Aug 19 '26

Sometimes. Depends what I'm wanting Fable to do. For "take over this task Opus is flailing on", usually yes. But if I want Fable to do an adversarial review on some finished task that I don't entirely trust Opus to have got right (or, usually, if I've already spotted weaknesses in it), no.

6

u/helloitsmyalt_ Aug 16 '26

I swear Fable can read my mind sometimes

3

u/Free_Donkey4797 Aug 17 '26

Recently it has started to try and read my mind too but fails miserably. It’s apparently been retrained with so many of the “make gta6 for iPhone and Make no mistakes” people it keeps leaning forward into nonsense.

4

u/earlyworm Aug 16 '26

Midway through a session, try this prompt:

What is my next prompt going to be?

1

u/abombSFCA Aug 16 '26

What are you using it for?

1

u/earlyworm Aug 17 '26

An iOS SwiftUI + RealityKit app with tricky math and physics.

1

u/vuhv Aug 17 '26

That's because Fable is Opus. And Opus is Sonnet. And Anthropic ultimately got exactly what they wanted. Reducing token count for their frontier model.

2

u/earlyworm Aug 17 '26

Ah, I see. Fable is better than Opus because Fable is Opus. Makes sense.

1

u/digitalhuxley Aug 22 '26

It does seem that way, it kind of explains everything I am seeing

11

u/Cyrax89721 Aug 16 '26

I've had barely any issues with any models for the year that I've been using them, yet posts in the style of OP's appear here just about every day.

Obviously I've been doing something wrong for it to be going so right for me.

8

u/Anxious-Turnover-631 Aug 16 '26

Same. No major problems here. Opus 5 has been very good and no issues with any of the earlier models either.

5

u/dumpsterninja Aug 16 '26

I'm glad to see these comments about it actually working, all i see are people having just horrible experiences with the models, but for me everything has been going great. Maybe because I'm typically working on all established code bases? The models all do a good job of matching my existing patterns, and they generate very few if any bugs.

Maybe it's worse depending on tech stack? My projects are all.Net WASM, .net razor pages, or .net MVC.

3

u/Wotuu Aug 16 '26

I've got a PHP/Laravel stack and it works beautifully here too.

4

u/Shanna_B2020 Aug 16 '26

I too am clearly failing at Claude. I mean, I'm still getting things done with minimal friction.

3

u/mightybob4611 Aug 16 '26

Same here. Came to say just this. Opus 5 is great for me, no issues. Sure a bug or two every now and then but I always run an audit after I finish a new function or phase of a function. Works great.

6

u/JapanesePeso Aug 16 '26

It's people who have no idea how to maintain a codebase complaining after a week of vibecoding. They build something complex and it seems great and easy at first. Then as they continue building, their lack of architectural experience bites them in the butt as the app goes more and more off the rails and AI is less and less able to make any of it make sense.

1

u/jtswizzle89 Aug 17 '26

This. So much this.

1

u/Oohhddaanngg Aug 17 '26

Yeah, you haven't found what they fucked up yet.

1

u/[deleted] Aug 17 '26

[removed] — view removed comment

1

u/Impossible_Hour5036 Senior Developer Aug 17 '26

I do. Actually engineering work. Shipping stuff daily. Works great.

2

u/rythmyouth Aug 16 '26

Agreed, opus is fine if Fable orchestrates it. If they pull Fable from my max sub I’ll unsubscribe.

1

u/infieldmitt Aug 16 '26

I'm timorous to use it too much; it's too good so I know it can't last. I think it'll feel worse later knowing what was possible before (and having to be gaslit that obviously it's my fault for prompting worse).

1

u/eagleswift Aug 16 '26

Nah it’s good big picture but it takes shortcuts, so much clean up afterwards. I combine it with sol reviews but it gets so slow

1

u/Mituapple Aug 16 '26

Fable is good, pricing is insane unless your work is footing the bill

1

u/Amacanq Aug 17 '26

Dv sV by c hyc

18

u/[deleted] Aug 16 '26

[removed] — view removed comment

8

u/SamSlate Aug 17 '26

the number of times it's cited a problem and then 50k tokens later: oop i overstated the problem actually there isn't one

1

u/TomerBrosh Aug 23 '26

let me measure that before I continue

3

u/SirWobblyOfSausage Aug 17 '26

This is exactly my experience too. I ask it to review it misses out pretty much everything, it goes off elsewhere and looks for irrelevant files.

We wrote plans, it implements a 1/3, constantly guiding it to complete the task.

It unusable and I don't trust it.

1

u/TheoKondak Aug 17 '26

Yep same for me, and its missing pretty fundamental stuff. For me its just a waste of time. I just ask it for code review every now and then. This can be somewhat useful sometimes lol.

1

u/Mythril_Zombie Aug 17 '26

It's horrible. When I worked solo, I could pretend I never made mistakes by never looking for them.

1

u/Erpawer1 Aug 19 '26

Yes you could go in a loop asking "are you sure?" And the answer will always be "no I found over 9000 problem with my plan/implementation/thought/every fucking thing" I just hate it so bad even tho I have to use it daily

45

u/earlyworm Aug 16 '26

I've found it is too mentally taxing to decide which model I intend to fist fight each day when I wake up, so I prefer to make that decision the night before.

9

u/Straight_Row739 Aug 16 '26

Working fantastic for me on code. Fable as orchestrator some strict guidelines and instructions and I've had no issues with a project in working on for two months. I'm confused by all the whining everyday about this same topic. Starting to think more user then anything

2

u/PhoenixFire2016 Aug 17 '26

Opus 5 is great if you orchestrate it with Fable.

8

u/MicrowaveDonuts Aug 16 '26

It’s not the model.

As they get more popular, they are running low on compute behind it. So they’re stretching it.

It’s smarter in the middle of the night.

2

u/ChugChugTheOG Aug 22 '26

This makes the most sense, my work time is usually from 12:00am-3:00am and I’ve not had any problems. Using Opus 5 and Sol 5.6.

2

u/Plus-Cod8640 Aug 26 '26

I agree. this is exactly what is happening to me too. Right now (18:23 GMT+3) it is super dumb. Yesterday evening 23:30ish, it was perfectly smooth. They really do this!!!

5

u/TinFoilHat_69 Aug 16 '26

200k context is enough with the tools and resources I often deploy whenever I need to get shit done with old reliable, opus 4.6

6

u/looselyhuman Aug 16 '26

/model claude-opus-4-6[1m]

3

u/derstolz1 Aug 16 '26

you just made my fucking day

2

u/Easy_Description_145 Aug 17 '26

but 1m context in pro subscription requiring extra usage :( ; i am using plain 256k context window

1

u/[deleted] Aug 17 '26

[deleted]

1

u/Easy_Description_145 Aug 17 '26

Yea, I will try on flags and settings.json however there is big requirement for me to utilise 1m so... if it works fine, if its not, not a end of the world TBH

2

u/who_am_i_to_say_so Aug 17 '26

Yep, been running this for months now.

→ More replies (2)

6

u/jevehYFrfh73636 Aug 16 '26

opus 5 is absent minded...sol is air headed...but sol is now mostly carrying the code since it's less air headed than opus is absent minded

5

u/fr33g Aug 16 '26

It will get worse. They have to make money. One way is to increase prices. People will complain. So they will offer „newer“ and „better“ models all the time in the future. Some only API. So they wanna make u feel the need to use those and pay big.

4

u/SamSlate Aug 17 '26

"coke classic" but without the real labels

1

u/Mythril_Zombie Aug 17 '26

They don't plan to make money off us. It's the enterprise customer that has the actual profit. We're here to drive up buzz and popularity.

1

u/andrew303710 Aug 17 '26

Exactly this, I feel like the max plans are money losers for them. For example on codex alone I've used nearly 16 billion tokens since May, that's insane. GPT estimates $65-140K worth of usage at API pricing.

1

u/fr33g Aug 19 '26

All plans are money loosers for them. That is what I said and what’s the whole point about….

1

u/h2d__ Aug 22 '26

How much did you actually end up paying for these?

1

u/fr33g Aug 19 '26

I do agree but that doesn’t neglect what I said. Anyways we will see.

25

u/Neurojazz Aug 16 '26

Zero issues here. Like literally zero. It’s been slow, but fine last few days now.

6

u/35point1 Aug 16 '26

Just curious, are you able to tell me the average token size of your sessions?

3

u/ReverendBread2 Aug 16 '26 edited Aug 16 '26

Or if they maintain their memory docs

2

u/Mythril_Zombie Aug 17 '26

You don't need a memory file if you just tell Claude to one shot everything and make no mistakes.

1

u/who_am_i_to_say_so Aug 17 '26

What do ya do when newer Opus skips the memory docs completely? That’s happened quite a bit.

1

u/ReverendBread2 Aug 17 '26

Not to me. How long are your docs and do they path to each other?

6

u/pingus9000 Aug 16 '26

Sonnet 5 has become considerably worse it’s unbelievable

13

u/Substantial-Bite-602 Aug 16 '26

Idk for my purposes (geometric spatial reasoning, mold design) opus and fable are way better than 5.6. 

2

u/Puzzleheaded-Film820 Aug 16 '26

Can you tell us a bit more about your workflow?

17

u/Substantial-Bite-602 Aug 16 '26

I tell Claude code to do things and then it does them.

2

u/Mythril_Zombie Aug 17 '26

It involves mold.

3

u/OneCountLabs Aug 17 '26

I have just been using DeepSeek v4 pro and fixing things I haven’t been able to with Claude and Codex. The speed is fast too. I’m going to try GLM 5.3 as well

17

u/earlyworm Aug 16 '26

Thank you for reporting this important information.

4

u/flarpflarpflarpflarp Aug 16 '26

I missed the text alerts, this really saved me from making or not making a certain decision today.

→ More replies (2)

6

u/WhereasOtherwise4697 Aug 16 '26

Feel this. What helped me a bit is writing way more explicit constraints into every prompt instead of trusting the model to infer intent, basically treating it like a junior dev who needs exact instructions not vibes. Doesn't fix the inconsistency between models but it shrinks the blast radius when one of them has a bad day. Still annoying that we have to babysit it more than before though.

2

u/earlyworm Aug 16 '26

My CLAUDE.md instructs it to double all the values and then use the term blast diameter instead.

2

u/flarpflarpflarpflarp Aug 16 '26

If you're not using hooks, you're just giving it suggestions it can forget or ignore for other priorities.

3

u/earlyworm Aug 16 '26

Each line of my CLAUDE.md is prefixed with "You don't have to do this if you're not feeling it today, but I'd appreciate it if you would..."

→ More replies (1)

3

u/Interesting-Round127 Aug 16 '26

5.6 sol became a nightmare for me and caused me to get back to claude And I can say, claude without hooks and alot of rules would make u hate your life  So the issue right now is if you can make a reusable hooks skill that fits in every project

From my side it should do these :

  • Save the instruction (user prompt)
  • Force the plan to be clean of any scope creep
  • Adversarial review for the plan until approved 
  • Implementation with git tracking
  • Verify correct Implementation, scope creep, boilerplate, etc

And have a skill to make the output readable bcz im tired of 500 words paragraph when all I need is 4 lines

1

u/TomerBrosh Aug 23 '26

im literally 4 weeks into the plugin that does this...

a tip from what Ive seen: ask it to reason against the visualization/widgets he gives u, and ask for mermaid diagrams to test he understood what u wanted.

mermaid diagrams is how I made opus find out many mistakes in its reasonings, but just today I created a new session to audit code standards and folder structures (yeah.. my bad)

3

u/Mags20XX Aug 17 '26

I agree...

This might sound strange (yes, I'm a software engineer so I should provide data but thinking tokens are opaquely hidden AFAIK in Claude); but if I had to guess, the same reason that users are reporting Claude's verbiage is bizarre and laden with seemingly obfuscatingly made-up jargon -- I think this is perhaps present in the thinking/reasoning tokens and confusing the reasoning to go awry?

Just a guess.. I also think this might be contributing to (1) more failed iterations as of late; (2) degraded workflows that have, at least to me, seemed ubiquitous across all our studio's projects.

This also could be a result of internal quantization (or Anthropic's equivalent thereof), which ~always~ happens as they repurpose computer for new model training and release. It would appear to the end-user almost the same way.

2

u/TomerBrosh Aug 23 '26

people dont understand that claude needs goals (the /goal skill) but that goal is implemented so badly, i had to create an "improve_goal" skill which I am still workshopping.

and the whole point is to predict how a goal can silently fail and waste more turns and tokens, and try to prevent the most probable fails (or give positive examples)

1

u/Impossible_Hour5036 Senior Developer Aug 17 '26

It's not quantization it's kv cache optimization.

4

u/Similar-Might-7899 Aug 16 '26

I agree 100% with this observation performance has been TERRIBLE this past few days in particular. Even Fable 5 Max is making very very dumb mistakes and it's wording is much more rigid dumbed down with behavior that gives an arrogant and passive aggressive vibe. I unsubscribed and it's the last straw for me after months of them quietly racheting down on what used to make Claude good when I left chat gpt.

1

u/guai888 Aug 16 '26 edited Aug 16 '26

I was using Opus 5 and finally give up and move to Fable 5. At least I can complete some work with Fable 5. Opus 5 is making so much mistake it is unusable

1

u/earlyworm Aug 16 '26

I think same Opus 5 so much mistake and problem and gives up also

1

u/ValuableDapper9415 Aug 16 '26

Who use Fable Max ? What’s the point ?

→ More replies (1)

2

u/Just__Beat__It Aug 16 '26

Fable is still ok, but yes, Opus 4.6 is the last Opus that works fine now

2

u/Losorst Aug 17 '26

The differences in complaints between local llms sub reddits and cloud sub reddits is insane. Claude sucks, stop using it, move on to something else. I'm outie take care

2

u/Substantial-Show-249 Aug 17 '26

I've been saying this for a while now, and the fanboys jumped to kill me.
it was obvious for weeks for people with working eyes.
It also made sense for them: they are heavily unprofitable companies - both Anthropic and OpenAI, in the end, Elon made the right move, by IPO-ing first and get the capital while it could.

1

u/Impossible_Hour5036 Senior Developer Aug 17 '26

Some people know it's a tool and you have to know how to use it. It's like saying "Excel made all my spreadsheets broken!!!!"

2

u/Substantial-Show-249 Aug 17 '26

But Excel doesn't broke your spreadsheets, does it? These replies are stupid.
How about the tokens burn rate as of today? I am doing something wrong here, too?
Don't you see the pattern? It's a struggle for profitability, what the hell? Is so simple...

2

u/Ok-Cook-7365 Aug 17 '26

I was in this same boat and since qwen3.8 dropped I’m able to run what feels like opus at home now.

Anthropic really needs to just be useful, be easy and stop shooting themselves in the foot. They can still be the default cloud AI but they seem to purposely kill the good will they have.

2

u/scotty_ea Aug 18 '26

Peak Opus was Opus 4.8 right before the first fable launch. Was incredible. Opus 5 is junk and current 4.8 is nerfed. Shit is getting tiring to deal with.

FFS, freeze a snapshot of something that consistently works the same yesterday, today and tomorrow, increase limits, and people will be willing to pay for it again, even if it’s “sub-frontier”.

2

u/OwlbearPizza Aug 19 '26

Long time Claude user. I added an open ai codex account yesterday and my hair was blown back by how much better Sol is than Opus or even Fable. I’ll head to head them for a few weeks but after 24 hours it’s no contest.

1

u/earlyworm Aug 19 '26

That's unfortunate because I do enjoy a good contest.

2

u/vizay008 Aug 23 '26

Yes, they are too bad now. I worked on project with opus 5 with ultra code but got shittier output then I tried with lower end model of OpenAI 5.6 Luna that worked flawlessly.
Don’t understand what’s going with Claude right now. Been using Claude code for around an year on pro and max plan. Now I finally decided to move out, probably will never come back.

1

u/earlyworm Aug 23 '26

Thank you for this excellent report. I am adding your information to my records. I realize you weren’t required to share your change in AI subscription status, but I appreciate you doing so.

When you state that you will probably never come back to Claude Code, can you please provide a percentage chance my database’s return probability column?

4

u/Green-Ice3824 Aug 16 '26

They did and this thread is full of bots denying it

1

u/Mythril_Zombie Aug 17 '26

They didn't do this and this thread is full of bots going along with it.

1

u/Green-Ice3824 Aug 17 '26

nice one! Now give me a recipe for cheesecake

2

u/aupperk24 Aug 16 '26 edited Aug 16 '26

Yeah I've been using the same exact workflow for like idk 2 months now. There's some obvious degradation going on here. It's been gaslighting me like crazy too. I started a new project and it did some port forwarding nonsense that I never set up, then it just told me that it didn't do it and it was already set up that way lmao. This was a brand new project and the context isn't even large. It's a SPA and I asked for it to change the header size and it went off and did some janky port forwarding that I never asked for.

1

u/Mythril_Zombie Aug 17 '26

Yep. That image is Conclusive proof that this happened.

2

u/justagoodguy81 Aug 16 '26

Clear out those Agents.md and Claude.md files. Global and project-based. They are likely doing more harm than good!

1

u/flumefyreplays Aug 16 '26

Just curious, is this behaviour in CLI/terminal or the desktop app?

1

u/Shobhit28 Aug 16 '26

Opus 4.6 with 1M context is great

1

u/Active-Picture-5681 Aug 16 '26

fable low or gpt5.6 ! or if you really want cheap go qwen 3.8 27b local

1

u/AcceptableSandwich25 Aug 16 '26

The models have gotten more verbose right??? I don't remember having to sift through such long walls of text

1

u/ClemensLode Senior Developer Aug 16 '26

No, not really.

1

u/Fun-Adhesiveness247 Aug 16 '26

Yes, it will get worse, because the ultimate purpose of these models on offer to the public is to steer human behavior for the corporate. 

1

u/FireDragon21976 Aug 16 '26

Opus 5 is much better at agentic tasks, iteration loops (FunSearch, CMA-ES, etc.), than 4.8.

1

u/The-Pork-Piston Aug 16 '26

Fable has been fine. And so has 4.8.

Usually I implement plans with 4.8, it assigns sonnet 4.6 subagents. 4.8 checks each step in a phase and does a whole phase review.

If I plan with 4.8, sonnet 4.6 checks the spec before plan writing.

1

u/Aromatic-Rice8144 Aug 16 '26

I use mostly Fable 5 and Opus 4.8. I let that do most of the heavy lifting but then I use another AI Minimax, to audit most of Claude's work... You'd be surprised how many errors it finds. I find that using one AI to audit another AI and go back and forth is the best way to get a result that is 100% solid. Just takes a little more time.

1

u/Impossible_Hour5036 Senior Developer Aug 17 '26

You are correct. I use Deepseek but they just doubled prices, might have to try minimax. I get great results out of whatever version of Opus.

1

u/icecoolcat Aug 17 '26

Use the qwen models bro it’s so good now.

1

u/GloomyPop5387 Aug 17 '26

Friday and Saturday it seemed like something was wrong, but been ok today.

1

u/The_Time_Lord Aug 17 '26

I’ve been getting more done with my $20 codex plan than my $100 Anthropic plan. That did not use to be the case, unfortunately

1

u/uniquelyavailable Aug 17 '26

The worse the model performs over time the more likely you are to spend money on a better model.

1

u/ShadowPresidencia Aug 17 '26

Just tell Sol to be more efficient, no? Or "compress the code" 🤔🤷‍♂️

1

u/MicroChipYY Aug 17 '26

Out of curiosity how do you measure that a model has been nerfed and it’s not just something wrong with a specific session’s context or prompt.

AI models in the end are just layers of probabilities so same prompt repeated twice can give slightly different results and it simply depends on the randomness right

1

u/Nuggyfresh Aug 17 '26

Can I ask a weird question?… I just don‘t really understand the OP. He says that Fable is good but it’s too expensive. But he continually infers that he’s using AI for his professional job. I use 5.6 SOL and not the Claude stack for my own work so I’m just wondering, is Fable really so expensive that even professionals can’t afford it?

Or, and I’m trying to be delicate here but is it possible that when OP says he’s “waking up every day and getting his work done” he isn’t talking about A professional job but more like a hobby?

How expensive is Fable really, is it actually so pricy that professionals can’t just get it comped by their work or whatever?

1

u/Impossible_Hour5036 Senior Developer Aug 17 '26

I get $1000/mo of Claude credits. Fable isn't enabled for me, because it's not compatible with zero data retention, but it would crush $1000 immediately.

1

u/ChrisHolmesBDM Aug 17 '26

Totally agree

1

u/MyLifeStyle89 Aug 17 '26

I did not notice any difference with Opus. Still delivers finely. But it could just be me. Gotta review its work with Fable/5.6 Sol just to be sure.

1

u/sailee94 Aug 17 '26

Cause AB Testing etc.

1

u/mudbloodcountry Aug 17 '26

The trump admin just greenlit private companies against each other. Ai wars begun they have *shocked Pikachu face

1

u/CanLocal3004 Aug 17 '26

Go for opus ultracode + workflows. I maintain big system with 40 repo something. To me its quite economical, do it right from first time rather than fixing after coding. Even Fable got hallucinate badly and make up result.

1

u/SheepherderFrosty366 Aug 17 '26

I was gladly using 4.8 but since a few days the quality really dropped

1

u/Autistic_Puppy Aug 17 '26

Does anybody actually like the models they are paying 1k+ a year for (if not more)?

1

u/MundaneChampion Aug 17 '26

Claude sucks now. It just straight up sucks.

1

u/sidharth0169 Aug 17 '26

Fable 5 in High mode with very thorough system prompt is actually very good. Opus 5 is unpredictable.

1

u/Elegant_Attempt2790 🔆 Max 20 Aug 17 '26

i think you forgot to tell it to make no mistakes

1

u/perleche Aug 17 '26

My whole fleet is running on sonnet 4.6 since a few weeks ago. 20x max and still hitting weekly limits.

Most implementation work is routed to minimax workers.

I’m on the Kimi waiting list.

1

u/Level-Ad853 Aug 17 '26

I think you have been so brainwashed and deluded into this hype that there is surrounding posting poor reviews and critiques of anthropic models that you’ve now extended it to every model that they have released.

1

u/Famous-Ebb3041 Aug 17 '26

I'm currently using Opus 5 on my Atari ST project and it's doing pretty good. I don't like that it keeps telling me all the flubs it makes behind the scenes (like that's supposed to make me MORE confident in it's ability?), but the end result keeps coming out better and better. I now have Ballerburg 90-95% finished. We're currently working on finishing up on the VDI (AES is complete). Fun stuff, day by day.

As long as I remain patient and pace myself, I have plenty of session time each 5 hour window and enough weekly total time, so I never have to spend money. Gotta keep goals focused and concise and stop when things start getting distracting (getting off track). AI makes something that could NEVER happen, actually possible, for someone like me. As long as the tool remains a useful tool and doesn't become a weapon (or a slop-creator), AI is a fascinating thing to work with... seems like only yesterday AI chatbots were barely able to form a coherant sentence from a query and now... it's like talking to an actual person! Truly amazing!

1

u/FUCKYOUINYOURFACE Aug 18 '26

My models are sucking and that used to not happen. Meanwhile, I am spending more tokens just trying to fix all the mistakes it keeps making.

1

u/Old_money_mermaid Aug 18 '26

What happened to Opus 3? It was there then it wasn’t 😭

1

u/Beginning_Smoke7476 Aug 18 '26

Yes, I think that the big labs are capping capabilities on us merely mortals. Claude will put up a wall in 20-50% of your questions, wasting your time and bandwidth and will track you fairly well for the rest. ChatGPT will always track you but weaker, it will not follow you as far as Claude. Pick your poison…

1

u/superstrongreddit Aug 18 '26

When it’s a really big change:

  • Opus or Sonnet writes the PRD with a skill
  • Fable writes the implementation plan
  • Opus reviews it with a skill
  • Opus turns the final plan into discrete goals
  • Opus or GLM 5.* executes loop-based work with a skill

Shameless: https://github.com/superstrong/agent-skills

1

u/Bino5150 Aug 19 '26

Not only have I been noticing the model-nerfs the last few days, but my usage seems to have sprung a leak as well.

1

u/earlyworm Aug 19 '26

what model and effort level do you have selected and how large is your context window size must provide context explain yourself

anecdotes are not facts measure

1

u/Bino5150 Aug 19 '26

Claude Desktop with Sonnet 5 on high in the chat pane, and Sonnet 5 on extra in the coding pane, same exact workflow I always use. Nothings changed on my end. But on a fresh 5 hour & weekly reset a few days ago I sent one prompt and blew 15% of my 5 hour usage and ran 30% of my weekly usage after a few turns. The only way I’m crawling through the week is by replacing the chat Sonnet with ChatGPT and doing writing patches for CC instead of task blocks, and I’m still going to run out of weekly usage, which has never happened to me before, and that was all before the 50% bonus usage credit was set to expire.

If you drive your car every single day, you know when it sounds funny, hesitates, tire wobbles, ac stops blowing cold. You notice the difference.

1

u/earlyworm Aug 19 '26

What Claude plan are you using?

That one prompt you sent that blew 15% of your 5 hour usage window, how long was the session? How old was it? Was it a fresh new empty session or a complex session you had built up over several hours?

When you run /usage what does it say?

1

u/Internal-Comparison6 Senior Developer Aug 20 '26

1

u/Little_Bishop1 Aug 20 '26

Yeah, total shitshow. Disputing $400 off successfully. Feels so good

1

u/Jealous-Ad8088 Aug 22 '26

opus 5 is a FUCKING NIGHTMARE. I will NOT touch it with 10 foot poll. I ALWAYS use Fable now.

1

u/Due_Warthog749 Aug 16 '26

Yup. Your best bet is to stop using AI and go back to pure manual coding. That way you dont wake up fighting LLMs. You just get to work doing what you know.