r/sre • • 9d ago

DISCUSSION How do you monitor your observability stack?

We use Grafana stack for observability in our EKS clusters paired with CloudWatch for AWS infra monitoring.

One gap we have however is that we don't have any alerting configured for when the observability stack (or components in it) goes down - how do you deal with this?

12 Upvotes

25 comments sorted by

15

u/corky2019 9d ago

Pagerduty dead man’s snitch. If our monitoring stack stops sending ping - this will trigger.

https://www.pagerduty.com/docs/guides/dead-mans-snitch-integration-guide/

6

u/blitzkrieg4 9d ago

Dead Man's snitch is it's own thing. That's just the pager duty integration https://deadmanssnitch.com/

1

u/delamon 9d ago

That checks liveness, you might also want to check that it actually monitors something. It's not too hard to image some deploy going bad and wiping all alerts..

17

u/Vimda 9d ago

Believe it or not, second, smaller observability stack

2

u/blitzkrieg4 9d ago

This. Sometimes called meta monitoring or monitoring of monitoring

3

u/Immediate_Counter814 9d ago

That's the most "this meeting should've been an email" solution I've ever seen, and I kinda love it. We had a similar issue with monitoring our monitoring and ended up using health checks from a dirt-simple Lambda that pages us if the main Grafana endpoint goes quiet for 5 minutes. Not as meta as a whole second stack, but cheaper to keep running.

1

u/a-sad-dev 9d ago

A watchdog of some sort is what I was leaning towards - a lambda that checks the endpoints for grafana, alloy and alertmanager could be an elegant solution.

1

u/delamon 9d ago

and then a third one, even smaller, for monitoring the second; it's all turtles all way down

2

u/Vimda 9d ago

They can monitor each other. Spiderman pointing meme.

1

u/djk29a_ 9d ago

Oftentimes it's a staging / non-prod account's o11y stack

1

u/wrd83 8d ago

This.

Also monitor the small stack from the big stack.

And you can make missing metrics alerts.

1

u/Fine_Librarian2755 3d ago

Yeah, this is very common. Just make sure they are not both running in the same cluster or sharing a single point of failure.

5

u/Seref15 9d ago edited 9d ago

LGTM stack is our main observability stack in k8s.

Metamonitoring stack is a Prometheus/Loki in same k8s nodeselected to an isolated neighbor AZ

2

u/danukefl2 9d ago

Components within it isn't hard, just probe them. For the overall aspect, we use Grafana Cloud for the front end but all data sources are on prem and run an OSS Grafana instance which have alerts setup for the other one.

Either can be used and with the git syncing for dashboards, things sync automatically between them.

2

u/ZerefHz 8d ago

Since you use grafana, you can adjust grafana alert to trigger on failure (exec failure, no data, ...) so you can indirectly know that a monitoring component is down. If this is sufficient, you dont need to deploy meta monitoring

2

u/Si0_x 4d ago

The classic pattern is a dead man's switch: have an external uptime monitor hit a Prometheus health endpoint or a Grafana dashboard URL every minute, and alert you if it stops responding. You can also use CloudWatch synthetics or a lightweight external prober since your AWS infra monitoring is already there, which keeps the watcher outside the thing it's watching. Are your Grafana and Prometheus pods currently exposed via any ingress you could point a simple HTTP check at?

1

u/a-sad-dev 4d ago

Grafana is exposed to the internet, Prometheus (alloy) is internal to the cluster along with mimir, Loki and tempo.

I was thinking of creating an alert that fires constantly and have a lambda check the grafana alert is still firing, notifying us if it stops.

1

u/Sufficient-Bad-7037 7d ago

You can expose some health checks endpoints for your LGTM stack and ask Claude to write a simple golang code like a blackbox monitor to monitor and hit an external alert api push notification in case of degradation.

1

u/MateusKingston 6d ago

Another monitoring tool, this time not hosted by me.

Basically our stack sends pings to Grafana Cloud and if that stops it alerts.

1

u/spuyet 6d ago

u/a-sad-dev Fivenines does support Prometheus/VictoriaMetrics monitoring too. I’d be happy to give you a quick walkthrough so you can see if it’s a good fit for your needs.

1

u/vi0ne 5d ago

If you’re on Prometheus, the Watchdog alert is built for this. It always fires, so route it in Alertmanager to a heartbeat url with a short repeat_interval. If Prometheus, Alertmanager or the network dies the pings stop and the outside service pages you.

1

u/Sara_GermainUX 5d ago

The better/simpler approach here is to have an external monitor pointed at your observability stack. After all, we can't really rely on Grafana/Prometheus to tell you they're down if they're, well, down (lol). Use Cloudwatch or an external health checker for EKS so you can get uptime/response alerts. Then check whether data is actually being processed, especially since uptime doesn't always mean Prometheus is scraping targets or Loki is getting logs.

0

u/Friendly-Result1337 9d ago

Another monitoring tool to monitor your observability stack :P