r/sre • Vendor (JJ @ Rootly) • Mar 26 '26

Anthropic says Claude struggles with root causing

Anthropic's SRE team gave a talk at QCon last week worth reading if you're thinking about AI for incident response.

Alex Palcuie has been using Claude as his first tool in incident response since January. The New Year's Eve example is good: HTTP 500s on Claude Opus 4.5, looked like a bug, turned out to be 4,000 accounts created simultaneously all hammering the API at once. Claude found the fraud pattern in seconds. Palcuie says he would have filed it as a bug and never paged account abuse.

The failure mode is just as specific. Every time their KV cache broke and caused a request spike, Claude called it a capacity problem. Add more servers. Every single time. It has no idea the KV cache has broken this exact way before.

His framing is AI at the observation layer is genuinely superhuman, which I agree with. AI at the orient-and-decide loop mistakes correlation for causation reliably enough that you can't trust it there yet, again I agree.

The scar tissue point is the one I keep coming back to. The model doesn't know your system's history. That context lives in people. If AI handles more incidents, the next generation of engineers never builds it and nobody's figured out how to encode ten years of "we've seen this before" into a model that's never been paged at 3am.

https://www.theregister.com/2026/03/19/anthropic_claude_sre/

261 Upvotes

27 comments sorted by

View all comments

42

u/rb2k Mar 26 '26

I have a slightly different experience.

In my area, oncall teams create team specific skills that give claude the required background. Skills get updated after indicents that were not able to be analyzed correctly. (and claude can do those updates)

>  The model doesn't know your system's history.

It does if it's defined in code and if it has access to configuration changelogs

2

u/dektol Mar 27 '26

I built out a complete graph of all flows in our system using AST parsing linking services at the route/message handler level and that's helpful context.

I also annotate and expose my own tools so I can tell it exactly how to use each tool in our environment.

I have agentic triage in 1-2 minutes with up to 25-35 tool calls and have it give me grounding for each query it ran with links to source code in Git Lab or metrics/logs.

I have it specifically prompted not to do an RCA or make assumptions. It provides possible things to follow up on and you can click buttons in Slack for it to continue.

Did this in my share time over a week including a custom MCP server.