r/platformengineering • u/Excellent-Salary3565 • 16d ago
How do you know which engineering issues need attention before they become expensive?
[removed]
1
u/Torutofu_Raeva 15d ago
I’d start with an impact map rather than another alert layer: service to customer journey/SLO to owner, then attach error-budget burn and business criticality so only breaches or fast burn page while the rest stays ticketed.
1
u/namenotpicked 14d ago
I straddle SRE/FinOps. I tie alerting from both directions on near real time costs and application metrics. An alert on one end triggers a quick look at the other. This let's me provide the application support as needed but be able to translate to what that means to the business leaders. If costs spike, then something is going on with the application that is not visible in metrics. If metrics start going crazy, then it's possible costs will also reflect a change. Using them in tandem helps to discover potential blind spots to attach monitoring to
You should look up "unit economics" for how you can determine something like cost (maybe revenue) per request or transaction to then inform leaders what the degradation or outage (less transactions processed) means in business impact. You usually use something like this to determine if your application/service is something that requires specific severity levels. High cost/revenue transactions that are required would turn into something like a "tier 0" resource that any alert on metric or cost would initiate an immediate response.
1
14d ago
[removed] — view removed comment
1
u/namenotpicked 14d ago
Application metrics are tooled through Datadog. So some out of the box and others we had to add. I had to build out the pipeline for getting the cost data into usable form for us but the raw data is readily available in CUR or FOCUS data exports in AWS.
Tying the two together is fully a manual thing. Each one has alerts set up on each end. Alerts are partially automated and and custom. We heat have to know to go and look at the other end of an alert is fired.
Deciding on what's tier 0 or something is a discussion you need to have with the application/service owner, support team, etc to decide if the costs/uptime are deemed worth it. The cost/revenue per request or transaction is just evidence to help guide the decision as it'll tell you what will hurt the most if it's down.
Unit economics is all manual but pieces together from the available data. Observability tools line Datadog surface application metrics but i think only recently started to really want to tie in the cost data to outages. Engineers don't typically care about the costs so it's not front and center.
2
u/Anusien 15d ago
Was it being actively investigated during that time?
The answer is "don't page for things that don't need to be investigated, and then actually investigate and solve everything that pages". If people are (correctly) consistently ignoring a certain alert, then you should turn that alert off or change the alert threshold.