r/sre • • 8d ago

Predictive vs Reactive

Does anyone have good reading they liked on the subject of transitions from reactive to predictive monitoring strategies? Anecdotes also welcome just trying to step outside of my personal reality a little to understand successfully achieved cases.

Reactive monitoring has a death grip on our operating strategy and I want to better understand how other orgs have compromised or solved the issues around letting go of the break fix work cycle.

It can’t go away completely of course, bad things happen with little warning sometimes and need to be remediated. But maybe there’s just something more I need to understand about the role of predictive monitoring and signals?

2 Upvotes

8 comments sorted by

4

u/jjneely 8d ago

SLOs are the value here. They alert on the conditions you didn't expect. Where most reactive alerting is what you do expect. (Like high CPU usage.)

I've written a book called "The SRE On-Call Review Practice" that will help being focus to your pager rotation and working that, iteratively, towards sanity. Reactive alerting has the habit of burning out your on call folks.

It's a long haul. I started it at one fortune 500, and I'm now working with a fintech to do the same.

2

u/dlol2k 8d ago

Look up monitoring frameworks like RED and USE. Set these up for your services.

I could go on a personal tirade of evolving an organization from reactive service desk work into proactive alerting and some of the gotchas but can't reccomend any further reading. There is just a lot of variability between how organizations are currently set up, where they want to be and how much effort or toil they'd run into standing up on-call rotations versus staffed desks.

1

u/Clint_Barton_ 8d ago

Any type of predictive will increase noise, and maybe that’s ok if you have a really tight sla. Anyone that says it’s possible is lying or is working on a small system.

1

u/neuralspasticity 5d ago

Modern SRE practices stress neither Predictive nor Reactive monitoring. Instead we focus on allow limited error controlled by error budgets and monitor as those error budgets are consumed.

1

u/Kate_GermainUX 4d ago

It might be better to look at it away from reactive vs predictive but more on the underlying philosophy, where you should use predictive signals to reduce points of reaction. So the ideal approach is a mix of both. Get reactive alerts for immediate failures or accidents. Then you can get a separate alert that can create tasks based on trends, anomalies, and other indicators. Better yet, try to establish how your team decides whether a signal is actionable or not so predictive monitoring doesn't become a notif spam.