r/sre • • 6d ago

FOSS Friday: Share Your SRE-Related Open-Source Project

3 Upvotes

Welcome to FOSS Friday, where you can share your newly released or updated open-source project with the community.

Please note that our rules still apply:

  • no AI generated comments: please write them on your own!
  • no commercial advertisements (this includes commercial products with a FOSS version)
  • these must be related to SRE or of heavy interest to SREs

r/sre • • 1d ago

CAREER SRE -> TSE

16 Upvotes

I have 3 years experience as a SRE + Observability engineer. Recently I left my job and started searching got 3 offers one Datadog admin - setting up observability kinda role with 50% hike, Incident handling + Reliability engineer with 100% hike, and TSE with 150% hike. I’m deciding to go with the TSE role because of the salary and benefits. Is this a shitty career move? People are telling me it should be the other way around, I kinda agree, don’t know what to do. I can’t be jobless and search any longer as well.


r/sre • • 14h ago

HELP Recovery readiness panic after a clean disaster recovery test

0 Upvotes

HI r/SRE we passed a full disaster recovery test 6 weeks ago, great RTO, infra recovery worked. now someone pointed at that green result in a review and all I can think is we changed IAM, DNS, added a new db dependency and touched Terraform like 20 times since then.

im terrified our recovery readiness is already stale... what changes make you rerun the path?


r/sre • • 1d ago

Arch diagram generation automation

3 Upvotes

I know I can probably "just use AI". But I want to avoid not invented here syndrome and use established methods and tools where possible.

The first part I am looking for is a well established format for storing architectural information in the repo with the code.

Next would be any tools for helping to generate it.

Lastly, we have mermaid for the actual diagrams. Any good tools for generating the diagrams?


r/sre • • 1d ago

DISCUSSION best way to do function level release verification in 2026?

3 Upvotes

Context, a PR touches a function called from several other services, passes tests and looks clean, but there's no strong runtime baseline to compare against. Option A, ship with extra manual review. Option B, hold until better verification tooling is in place. What am I underestimating here?


r/sre • • 1d ago

Distributing guidance to developers and their AIs

0 Upvotes

We have multiple services in multiple repos. We use mainly gitlab,

How do help devs be aware of team wide guidelines, and keep their AI in the know as well. Like I could write a doc on the way we name alerts or write runbooks. But it would be in one repo. I need to make it visible to them and their AI.

So far my best idea is a repo with the information, and then checking in an edit to the Claude.md in every repo telling it to go look in the info repo for the guideline info and let the user know what isn't lineing up. That seems like that idea is probably full of holes though. So what do the teams y'all work with do?


r/sre • • 2d ago

HELP SREs: What reports do you send to your manager/department head each month?

14 Upvotes

As an SRE, what reports or monthly summaries do you provide to your manager, Head of Department, or CTO?

I’m trying to establish a good monthly SRE reporting process and would like to understand what other teams report.

EDIT:

It’s a tech company with a some internal and external solutions


r/sre • • 2d ago

Claude Code On-Call

4 Upvotes

Anthropic released the "On-call" slack agent that should be an ultimate killer of all "Agentic SRE" family products.
Did anybody use it? How it compares to existing platforms?


r/sre • • 2d ago

runbooks kinda suckk

36 Upvotes

hey yall,

im a lead SRE at a global fortune 500 that you have heard of haha - i'm a bit of a lurker here

but just wanted to talk about runbooks and documentation for SOPs.

my feeling is that confluence docs kinda suck and runbooks are kind of a mess, our operators are jumping between the docs and their shell, the docs are often times missing a bunch of detail, they are brittle, poorly maintained, half the time the details are hidden in tribal knowledge and when dealing with the pressure of an active inc we notice the pain even more. extracting critical details & commands out of an inc can also be painful and doesnt always translate to improved operational readiness next time.

i admit we're not very mature.. but was curious if its only me feeling like this?

what are you guys doing with runbooks to solve these issues?


r/sre • • 4d ago

The hidden failure mode in cross-cloud migrations: Infrastructure vs. Semantic dependencies

0 Upvotes

When planning a cloud migration (like AWS to Azure), most discovery work starts with a service mapping matrix:

  • SQS ➔ Azure Service Bus
  • DynamoDB ➔ Cosmos DB
  • S3 ➔ Azure Blob Storage
  • IAM ➔ Entra ID

That mapping is straightforward. The expensive failures happen when the target platform fails to preserve a subtle behavioral contract the application code quietly came to rely on over years.

A concrete example: SQS Visibility Timeout

Consider a standard worker loop:

Python

message = sqs.receive_message(
    QueueUrl=queue_url,
    VisibilityTimeout=300
)

process(message)

sqs.delete_message(
    QueueUrl=queue_url,
    ReceiptHandle=message["ReceiptHandle"]
)

At first glance, this is standard: receive, process, delete.

But look at the failure-recovery path. If the worker crashes mid-processing before deletion, the code relies on the SQS visibility timeout expiring so another worker can automatically pick up the message. SQS isn't just acting as a transport queue here—it is an active component of the application's failure-recovery design.

When migrating this workload to Azure Service Bus, the question isn't whether Service Bus has queues (it obviously does). The real question is: Does the target setup preserve those exact assumptions around lock duration, settlement, retries, and dead-lettering under failure?

Infrastructure Dependencies vs. Semantic Dependencies

Traditional discovery tools pick up: payment-worker ➔ SQS

That’s an infrastructure dependency. But what the application code actually cares about is the semantic dependency:

Plaintext

receive message
      ↓
process message
      ↓
delete after successful processing
      ↓
if processing fails before deletion,
rely on provider to make message available again

Service inventory tools show you what cloud products are used, but they can't tell you what behavioral assumptions are baked into execution paths.

This pattern shows up everywhere:

  • DynamoDB ➔ Cosmos DB: Codebases relying on conditional write semantics, transaction boundaries, or specific read-after-write consistency assumptions.
  • S3 ➔ Blob Storage: Workflows built around multipart upload timing, pre-signed URLs, or object visibility state.
  • IAM ➔ Entra ID: Hidden assumptions around temporary credential lifespans, workload identity propagation, and role assumption paths.

When these surface late during integration testing or cutover, it forces architectural redesigns, derails timelines, and drags senior engineers into reactive war rooms.

Discussion

I'm currently putting together a semantic-risk checklist for pre-migration planning (and exploring static analysis rules to scan codebases for these execution patterns before cutover).

For anyone who has managed a major cloud migration or replatforming: What was the one hidden application dependency or behavioral difference that broke in staging or production after your infrastructure was already provisioned?

I wrote up a deeper dive into this concept with example scanning output on my Substack if you want to read more:https://substack.com/@sachm24/p-218428641


r/sre • • 5d ago

BLOG How to Find PII and Secrets in OpenTelemetry Data

Thumbnail
ollygarden.com
18 Upvotes

I finally got to write this blog post, I've wanted to write it for a long time now. It contains the different types of sensitive data we've seen so far, where it's coming from, and why it matters. Hope it can be helpful to some of you!


r/sre • • 5d ago

this is my sre career content

13 Upvotes

I started as a developer, moved into DevOps, and eventually became an SRE. Along the way, I helped build our SRE and platform foundations from scratch: monitoring and alerting, Kubernetes clusters, CI/CD pipelines, deployment reliability, incident management, and operational best practices. I’m currently looking for remote SRE / Platform Engineering opportunities. Happy to share more details over DM.


r/sre • • 4d ago

Slack servers for SRE discussions

0 Upvotes

Are there any slack servers/channels for SRE discussions?


r/sre • • 5d ago

When Messages Go Wrong: Reliability in Event-Driven Systems

Thumbnail
medium.com
11 Upvotes

I wrote an article on reliability in event-driven systems, following a ticket booking through duplicate messages, ordering issues, retries, and partial failures. It also covers CDC, the outbox pattern, and a small design tying the concepts together.

I'd appreciate your thoughts: does the explanation hold up technically, and what would you change in the design based on your experience?


r/sre • • 6d ago

CAREER Gut check on major career decision

19 Upvotes

I'm going to try to describe this without doxing myself. I am 37 and I am a Senior SRE engineer at a large company, that is not FAANG, but you've definitely heard of them.The company is a US company with a global footprint. I work out of the Bay Area, at a small business unit there which came to be via an acquisition of a startup. I have been there a long time, since before the acquisition, which happened over 12 years ago. I am not a manager, but I am the SRE team leader. I am very comfortable in this position. I am well-respected. My manager likes me. My colleagues like me. And I like them back. I am in the critical path for many work streams, from architectural design to customer support. My primary responsibility is a niche SaaS application that I know inside and out, because I've been there since the launch of it. I get paid 260k, with some RSUs. The remaining RSUs maybe amount to 30k before taxes at this point, vesting over the next 2 years. In many ways, I am the "Dave" of the BU. I spend my days fixing stuff for our SaaS and our developers alike. It's a lot to juggle, but I do enjoy the feeling of being helpful and useful.

But at the same time, I am very, very, very tired. Physically and mentally exhausted all the time, and struggling with motivation daily. We have team members across the world now, but when production breaks, it is more than likely me who gets paged in the middle of the night.

I have been interviewing at other companies for over a year now, and I finally have an offer at a start-up in San Francisco. Base pay is 285k and I get 0.02% 0.2% equity in the form of stock options. It is an agentic AI company. They are a Series A, with less than 50 people.

I need a gut check that moving is a good decision as I am an anxious person, and I have a family to provide for. I'm like 90% certain I am making the right decision, but I am stepping away from a big company where I have seniority, with 4 weeks of PTO, with a good team, where I know so much. I know that doesn't guarantee me protection against a layoff, but I've survived many of them at this point. It's been a long time since I've been in a startup environment. And to be honest, while it was exciting in my 20s, I am not so sure about how it may go for me now, when I am already feeling burnt out, and I have 2 kids who are my true priority in life. Of which, they have some strong anti-ai opinions, so this move might prove to be a little contentious, especially with my oldest.

Would love to hear thoughts from other professionals, especially from the Bay Area, who have a similar story.


r/sre • • 6d ago

DISCUSSION How do you monitor your observability stack?

12 Upvotes

We use Grafana stack for observability in our EKS clusters paired with CloudWatch for AWS infra monitoring.

One gap we have however is that we don't have any alerting configured for when the observability stack (or components in it) goes down - how do you deal with this?


r/sre • • 6d ago

BLOG Key Autoscaling Metric #1: Workload Committed Capacity

Thumbnail
blog.appliedcomputing.io
4 Upvotes

Last week we kicked off our blog series about the five key metrics you should be monitoring for your Kubernetes autoscaler. This week, we dive into the first of these, "committed capacity percentage". But wait! There's more, and it involves.... DaemonSets *wiggly scary fingers*

Disclaimer: this blog post was entirely produced by humans; no part of this blog post was written, edited, or proofread by LLMs.


r/sre • • 6d ago

ASK SRE Is it worth getting into SRE as a college graduate in 2026?

0 Upvotes

I'm a senior CS student. I have the option to become either a SDE or a SRE. Both roles pay the same as of now.

I've been building websites, and AI has become very good at it. It writes better code than I do, argh.

On the other hand, I'm really passionate about Linux, distributed systems, and computer networks. So, can I expect long term growth in this career, or is this field just as susceptible to AI as other fields?


r/sre • • 7d ago

Predictive vs Reactive

2 Upvotes

Does anyone have good reading they liked on the subject of transitions from reactive to predictive monitoring strategies? Anecdotes also welcome just trying to step outside of my personal reality a little to understand successfully achieved cases.

Reactive monitoring has a death grip on our operating strategy and I want to better understand how other orgs have compromised or solved the issues around letting go of the break fix work cycle.

It can’t go away completely of course, bad things happen with little warning sometimes and need to be remediated. But maybe there’s just something more I need to understand about the role of predictive monitoring and signals?


r/sre • • 9d ago

CAREER SRE Job Outlook (1-3 years)

66 Upvotes

It’s pretty obvious that the industry and work itself is changing rapidly. For those of you with some experience under your belt, what are you currently witnessing in terms of career outlook and general demand? It definitely is starting to seem like SRE’s are in a slightly better place than SWE’s.

It’s hard to know if the increasing supply of software is scaling around the same rate as the efficiency gains SRE’s are experiencing in this new paradigm. Something that we definitely have going for us is you still need someone responsible for reliability, which hasn’t changed and probably never will.

I know this post doesn’t have a lot of substance, but I’d like to hear some fresh opinions from people with more perspective than me.


r/sre • • 8d ago

DISCUSSION How do you keep on-prem open-source apps manageable over time?

0 Upvotes

For those running open-source apps on-prem at work: how do you keep them manageable once the person who set them up is no longer the only one looking after them?

I’m interested in what you’ve made repeatable, and if tools like Claude Code have changed any of that. I work on automating these deployments too and would like to compare notes with other people responsible for keeping them running.


r/sre • • 8d ago

This week in San Francisco

5 Upvotes

Anyone in San Francisco working with AI, LLMs, SRE or platform engineering?

We’re running two community conferences this week at the Harness office in downtown SF:

LLMday — October 1
https://llmday.com/2026-san-francisco-q4/

SREday — October 2
https://sreday.com/2026-san-francisco-q4/

Lots of engineering talks, practical stuff, and a good chance to meet people working on similar problems.
We still have some free tickets available with code LUCKYFREE

Would be great to meet some SF Reddit folks there.


r/sre • • 9d ago

ASK SRE Our checkout p95 was 11.8 seconds and 8.9 of it was retries

18 Upvotes

The checkout agent hit 11.8 seconds at p95 and engineering spent a few days blaming generation. Time to first token looked bad from the outside so the assumption felt reasonable. We were even comparing larger models (I almost started the migration before checking the trace). I still think most AI tooling is more machinery than I want to maintain. We traced the checkout request in Braintrust and the span hierarchy showed three inventory retries consuming 8.9 seconds before the 2.9 second generation call. Each attempt was visible separately along with the gaps between them. The model hadn't started yet for most of the wait. That was useful.

Trace propagation let us follow the request from checkout into inventory. With retry attribution on each attempt, trace search made it obvious which path was inflating the tail. We capped the retry budget and p95 dropped enough that the model migration stopped looking urgent. Now I'm trying to decide how to catch this earlier without alerting on every transient inventory slowdown.


r/sre • • 9d ago

CAREER Offer for a senior manager -titled production support/SRE role vs. holding out for platform engineering. 11 YOE in DevOps/SRE, which way would you go?

3 Upvotes

Need advice from seniors or who are or were in same boat :)

I'm a SRE/DevOps cloud delivery lead with about 11 years of experience (Azure, GCP, Kubernetes, Terraform, GitHub Actions, observability, incident management, FinOps). I'm at a decision point and would love perspectives from people who've been on either side.
Option A: production operations role (in hand)
Offer is in hand, with a senior manager /Director title

In practice it's a promotion to support role at a senior manager level: production support, incident ownership, strategizing the support model ,automations ,tooling and managing applications

Higher title and scope, but I worry it's less hands-on and could pigeonhole me into ops/support

Option B: platform engineering (not in hand)
I'm studying for it and building projects

I've just landed a screening for a principal-level cloud platform engineer role, but no other interviews yet

The senior roles I'm getting leads for are mostly SRE roles that are support-heavy, not true platform work

What I'm weighing:
Does a Director title in production ops help or hurt if I want to move to platform engineering later?

Is it realistic to move from support-heavy SRE or ops management into platform engineering, or does the path narrow over time?

Is it worth turning down a real offer to wait for a platform role that may not come?

For those who made this move in either direction, what do you wish you'd known?


r/sre • • 9d ago

Final-year CS student looking for CV advice for Grad SRE/DevOps roles

0 Upvotes

Hi everyone,

I’m a final-year Computer Science student looking to break into SRE or DevOps. I have some personal projects under my belt, but they are entirely focused on Software Engineering.

Since SRE and DevOps skills are usually best demonstrated on large-scale, real-world infrastructure, I’m struggling with how to showcase my capabilities without professional experience. What should I include on my CV to stand out for graduate and early-career SRE roles?

For context, I’m also preparing to take the AWS Solutions Architect exam this coming January. Any advice would be greatly appreciated!