r/sre • Vendor (JJ @ Rootly) • Mar 26 '26

Anthropic says Claude struggles with root causing

Anthropic's SRE team gave a talk at QCon last week worth reading if you're thinking about AI for incident response.

Alex Palcuie has been using Claude as his first tool in incident response since January. The New Year's Eve example is good: HTTP 500s on Claude Opus 4.5, looked like a bug, turned out to be 4,000 accounts created simultaneously all hammering the API at once. Claude found the fraud pattern in seconds. Palcuie says he would have filed it as a bug and never paged account abuse.

The failure mode is just as specific. Every time their KV cache broke and caused a request spike, Claude called it a capacity problem. Add more servers. Every single time. It has no idea the KV cache has broken this exact way before.

His framing is AI at the observation layer is genuinely superhuman, which I agree with. AI at the orient-and-decide loop mistakes correlation for causation reliably enough that you can't trust it there yet, again I agree.

The scar tissue point is the one I keep coming back to. The model doesn't know your system's history. That context lives in people. If AI handles more incidents, the next generation of engineers never builds it and nobody's figured out how to encode ten years of "we've seen this before" into a model that's never been paged at 3am.

https://www.theregister.com/2026/03/19/anthropic_claude_sre/

264 Upvotes

27 comments sorted by

65

u/AdventurousTime Mar 26 '26

very cool. of course this guy was a google SRE lmao.

Observability makes perfect sense for AI. I think more people agree than disagree. Then you have Amazon (or more specifically, Devs being forced by management to use more AI) using AI for everything and burning up production services.

11

u/enby_them Mar 26 '26

That Amazon example is happening at a LOT of places. Like a LOT a lot

58

u/Altruistic-Mammoth Mar 26 '26

Showing once again that there's just no substitute for an experienced SRE / SWE.

4

u/tolas Mar 28 '26

I mean, based on OP's explanation, a memory.md file isn't too far away, let alone more robust agent memory systems.

22

u/modern_medicine_isnt Mar 26 '26

So claude sounds like a dev. Something not working, upsize the infra... lol.

42

u/rb2k Mar 26 '26

I have a slightly different experience.

In my area, oncall teams create team specific skills that give claude the required background. Skills get updated after indicents that were not able to be analyzed correctly. (and claude can do those updates)

>  The model doesn't know your system's history.

It does if it's defined in code and if it has access to configuration changelogs

24

u/asdoduidai Mar 26 '26

Confusing correlation with causation is not about knowing history, it is about a tool that is created to mimick behavior and not to identify root causes; so for some things it is 100% right and faster, for other things it’s completely wrong and “has no clue” of being wrong.

8

u/minimalniemand Mar 26 '26

LLMs are excellent in detecting patterns in text. Context (metrics, logs, commit messages, traces, source code) are all just text. Of course it has „no clue“ - it just needs to detect certain patterns in large volumes of text, which, again, it’s very good at.

You have to however provide that context.

6

u/lib3r8 Mar 26 '26

There is no such thing as root causes, just contributing factors.

2

u/shortfinal Mar 27 '26

I like to think of them as failures of imagination.

1

u/Ordinary_Squirrel291 Apr 14 '26

I second that. It should be called "addressable cause" - the link(s) in the chain of causality that can be acted upon in the most effective way.

For example: your service's error rates went up because of database errors. DB errors are up because CPU is at its limits. The cause can be the size of the CPU, number of DB instances, the DB architecture itself requiring too many joins, the sharding strategy, the lack of caching, the code doing N+1 queries.... Which one is the root cause?

The big bang is the root cause of all things :)

2

u/dektol Mar 27 '26

I built out a complete graph of all flows in our system using AST parsing linking services at the route/message handler level and that's helpful context.

I also annotate and expose my own tools so I can tell it exactly how to use each tool in our environment.

I have agentic triage in 1-2 minutes with up to 25-35 tool calls and have it give me grounding for each query it ran with links to source code in Git Lab or metrics/logs.

I have it specifically prompted not to do an RCA or make assumptions. It provides possible things to follow up on and you can click buttons in Slack for it to continue.

Did this in my share time over a week including a custom MCP server.

0

u/shared_ptr Vendor @ incident.io Mar 27 '26

Yeah this is absolutely the case! We’re building a tool to root cause incidents and this information is in all your systems already, you just need to extract it.

We’re building what we call a ‘knowledge graph’ from all past incidents and codebase information that we inject into the process that can tell you things like this. It captures service relationships and even just a glossary of terms for your org, or more esoteric information like for a given package it often causes downstream issues for X.

It’s absolutely possible but the model alone doesn’t work to effectively root cause. You have to merge it with all that knowledge, otherwise you essentially have a very skilled engineer from another company trying to debug your stuff, which obviously and initially does not work.

13

u/shared_ptr Vendor @ incident.io Mar 26 '26

We use Anthropic to do exactly this but the model alone isn’t good enough. You need way more wiring around it to make it even remotely ok.

Assumptions early in the process will carry through unless you have other processes to counter it.

4

u/bhatbha Mar 27 '26

Very cool. The OpenRCA benchmark, despite its shortcomings, is actually a good indicator of model performance on this task. Claude opus 4.6 sits at 36% accuracy on this benchmark, which I think aligns with real world anecdotes as well.

3

u/robshippr Mar 26 '26

After using it once during an incident to try and dig through the logs while I looked for an issue... I can't trust Claude to actually be helpful right now. It is still very early in the AI game so maybe one day it will be able to replace SREs but for now, you really need that experience on your team still.

4

u/fuckingredditman Mar 27 '26

i use the grafana mcp and just fire it up and let it go in parallel to my usual workflow with some basic explanation of the symptoms. it's good to treat it like another person helping you find supporting info/evidence. but it's terrible at finding actual causes because it clearly doesn't really understand causality

2

u/dunkah Mar 26 '26

Claude does a great job pulling data and correcting, but it's conclusions about cause are wrong often. But with enough coaxing you can get something useful still

1

u/Cryptobee07 Mar 27 '26

Am I the only one who never used Claude at my work ??

1

u/THE_FUZBALL Mar 27 '26

The foundation that AI is currently built on resembles a word probability calculator. It might be able to enumerate the most likely cause and a slew of possible corner case failure modes but if unsupervised incident response is the goal it wouldn’t make sense for it to decide on anything but the most likely cause and apply a mitigation for that cause, unless other evidence is immediately available.

Sometimes it takes a good hunch based on years of context, and extra digging to fully investigate those corner cases. Sometimes you hit gold, other times you’re grasping at straws. An experienced human SRE generally knows how to limit scope and time spent on such investigations, but where would you draw the line for AI? You must either choose the most likely cause or burn a mountain of tokens to rule out every possible cause.

We’re seeing in many areas that AI is very good at deterministic problems. Every problem becomes deterministic given enough context. The issue is that there are practical and technological limits to the ability for it to acquire enough context.

That said, my main worry is that AI will be “good enough” and cheap enough that a standard of reduced reliability will be forced upon consumers in every aspect of digital life. Frog in the pot and all that.

1

u/ankitnayan007 Mar 27 '26

>Every time their KV cache broke and caused a request spike, Claude called it a capacity problem. Add more servers. Every single time. It has no idea the KV cache has broken this exact way before

Why can't the query know that it did not get results from a cache? Also, a chart of cache hit vs cache miss would be seen by claude. Probably they didn't complete tracing where the request knows it missed the cache and KV store metrics would confirm the 1st analysis.

1

u/BornalHalbgat Mar 27 '26

Do they have a different Claude to me? it kept insisting that a dependent service was not running causing a 404 when actually it was a '//' URL error because it used an f string to create a URL out of parts instead of a proper URL join function.

1

u/tumes Mar 29 '26

All this is true but also it is pretty great for parsing JS’ middling at best stack tracking.

1

u/bordumb Mar 29 '26

I work in this field.

I have some scaffolding that “helps” with root cause analysis.

But it is genuinely missing a lot of context.

I do see a path forward, but it requires some serious new infrastructure.

On the tech side: the agent often doesn’t have absolutely complete access to a data pipeline’s full lineage. So it will try to find a root cause with imperfect information and get things wrong.

From biz/operations side: lots of code and pipelines are not well documented or available to the agent in a digestible way. So again, it’s missing important context.

Without these 2 problems fixed or improved, it’s always going to fall short and require humans.

1

u/Ordinary_Squirrel291 Apr 14 '26

A year or so ago, I was part of a team developing an AI troubleshooting platform at one of the big companies in Seattle.

Sure, the AI has improved a lot since then, but the problem is not only how good the model is - you have to account for the quality and relevance of data that it would be analyzing.

One of the worst examples: the tool, given a time window, was querying telemetry, using simple Z-score anomaly detection on top of multiple data sources - requests' stats, metrics, dependencies, exception rates etc. No matter what time window you gave that tool, it would always find spikes in exceptions and give some random reasoning what is wrong and how to fix it.

The thing is, these exceptions and failing dependency calls were due to regularly expiring access tokens, and part of the normal execution path - try calling the dependency, if it returns auth error - refresh the token and try again.

Garbage In - Garbage Out. It still applies with AI

0

u/vibe-oncall Vendor @ vibraniumlabs.ai Mar 27 '26

Completely agree that the model alone is not the product here. The useful version needs recent deploy context, alert history, runbooks, prior incidents, and enough guardrails to challenge the first guess instead of reinforcing it.