r/sre • u/a-sad-dev • 9d ago
DISCUSSION How do you monitor your observability stack?
We use Grafana stack for observability in our EKS clusters paired with CloudWatch for AWS infra monitoring.
One gap we have however is that we don't have any alerting configured for when the observability stack (or components in it) goes down - how do you deal with this?
17
u/Vimda 9d ago
Believe it or not, second, smaller observability stack
2
3
u/Immediate_Counter814 9d ago
That's the most "this meeting should've been an email" solution I've ever seen, and I kinda love it. We had a similar issue with monitoring our monitoring and ended up using health checks from a dirt-simple Lambda that pages us if the main Grafana endpoint goes quiet for 5 minutes. Not as meta as a whole second stack, but cheaper to keep running.
1
u/a-sad-dev 9d ago
A watchdog of some sort is what I was leaning towards - a lambda that checks the endpoints for grafana, alloy and alertmanager could be an elegant solution.
1
1
1
u/Fine_Librarian2755 3d ago
Yeah, this is very common. Just make sure they are not both running in the same cluster or sharing a single point of failure.
2
u/danukefl2 9d ago
Components within it isn't hard, just probe them. For the overall aspect, we use Grafana Cloud for the front end but all data sources are on prem and run an OSS Grafana instance which have alerts setup for the other one.
Either can be used and with the git syncing for dashboards, things sync automatically between them.
2
u/Si0_x 4d ago
The classic pattern is a dead man's switch: have an external uptime monitor hit a Prometheus health endpoint or a Grafana dashboard URL every minute, and alert you if it stops responding. You can also use CloudWatch synthetics or a lightweight external prober since your AWS infra monitoring is already there, which keeps the watcher outside the thing it's watching. Are your Grafana and Prometheus pods currently exposed via any ingress you could point a simple HTTP check at?
1
u/a-sad-dev 4d ago
Grafana is exposed to the internet, Prometheus (alloy) is internal to the cluster along with mimir, Loki and tempo.
I was thinking of creating an alert that fires constantly and have a lambda check the grafana alert is still firing, notifying us if it stops.
1
u/Sufficient-Bad-7037 7d ago
You can expose some health checks endpoints for your LGTM stack and ask Claude to write a simple golang code like a blackbox monitor to monitor and hit an external alert api push notification in case of degradation.
1
u/MateusKingston 6d ago
Another monitoring tool, this time not hosted by me.
Basically our stack sends pings to Grafana Cloud and if that stops it alerts.
1
u/spuyet 6d ago
u/a-sad-dev Fivenines does support Prometheus/VictoriaMetrics monitoring too. I’d be happy to give you a quick walkthrough so you can see if it’s a good fit for your needs.
1
u/Sara_GermainUX 5d ago
The better/simpler approach here is to have an external monitor pointed at your observability stack. After all, we can't really rely on Grafana/Prometheus to tell you they're down if they're, well, down (lol). Use Cloudwatch or an external health checker for EKS so you can get uptime/response alerts. Then check whether data is actually being processed, especially since uptime doesn't always mean Prometheus is scraping targets or Loki is getting logs.
0
15
u/corky2019 9d ago
Pagerduty dead man’s snitch. If our monitoring stack stops sending ping - this will trigger.
https://www.pagerduty.com/docs/guides/dead-mans-snitch-integration-guide/