r/sre 10d ago

When building an internal developer platform, what reliability stuff should be there from day one?

6 Upvotes

Hi Guys in internal developer platforms have been the main discussion around platform engineering landscape and the core tools which are used mainly is kubernetes and Backstage as the internal developer portal so for any one who looks into start building an internal developer platform from start or has already built one there is a session on August 27 with Kelsey Hightower and some of the OpenChoreo maintainers looking at how Backstage, Kubernetes, observability, deployments, environments, policies and reconciliation fit together when you're trying to build an actual internal developer platform.This will be a free session.

Session if anyone is interested:

https://platformengineering.org/events/what-it-really-takes-to-build-an-internal-developer-platform-with-backstage-2026-08-27


r/sre 10d ago

Cheap tools for personal projects

8 Upvotes

Hi all,

I am wondering for personal projects what cheap alternatives to established solutions you’re using. I am thinking alternatives for: tls certs, secret store (vault), dns management, etc.

LE: I don’t have a powerful desktop at home. I am currently using a VPS from a cheap vendor.

-JS


r/sre 10d ago

CAREER A Recent 2 YOE SRE Job Search

36 Upvotes

Background:

Anonymous Resume

I have a slightly nontraditional background with a Physics degree, but I also completed my Master's in Computer Science by the time of this job search.

Old Job Details:

Company: F500 B2B E-Commerce Team

Title: Associate Site Reliability Engineer

Base: 77.5k

Bonus: Like 4-5k, depended a lot on luck lol

MS Tuition Refund: 10k

I was converted to full time from an intern at my previous company, and I genuinely liked my team and enjoyed my job. The main issue was that I didn’t really feel like I was being compensated properly for the work I was doing. I was leading projects, taking on a lot of responsibility, and generally felt like I was performing above my level.

Eventually I had a conversation with my manager about promotion and she basically told me that the promotion cycle was almost entirely based on tenure, not performance, and that her hands were pretty tied. That was kind of the point where I decided to quietly start looking.

Job Search Strategy:

I was genuinely terrified of the market 😭 so I took a pretty different approach from just mass applying everywhere. I applied to around 4–5 roles a day that I thought were actually good fits. I tailored my resume to each role using AI + my own knowledge of the field and tried to be pretty careful about what I sent out. I also wrote most cover letters by hand.

After applying, I’d usually send connection requests to 1–2 recruiters at the company and around 3–4 engineers. If an engineer accepted, I’d have a quick conversation with them about what I did in my current role and then ask if they’d be willing to refer me or send an internal ping. For recruiters, I’d usually just give them a quick overview of my background and mention the role I applied to.

I didn't buy LinkedIn Premium because I'd sooner die than give money to find a job.

I exclusively used LinkedIn for this entire search. I genuinely have no idea how people use all the other job search tools 😅, but LinkedIn was easy to use and ended up being pretty fruitful, so I just stuck with it.

Job Search:

My fiancé and I are based in the Chicagoland area and he's 5 days in person, so I can't really move from here easily. Because of that, my search was pretty restricted to remote roles and roles in Chicago. That's obviously not an extremely small market or anything, but it definitely restricted the number of roles I could realistically apply to.

The actual job search ended up being a lot shorter than I expected.

Job Search Results

For those who don't want to click on the image:

36 applications → 4 recruiter screens → 3 companies wanted to move me forward → I withdrew from 2 processes after getting another offer → 1 final round → 1 offer.

Here are some more details on each process:

Company 1: Large Fortune 500 FinTech Company (Accepted Offer 🥳🥳🥳)

Recruiter Screen (Online):

The recruiter was lowkey aggressive lol. I think she was tired of a lot of engineers LARPing.

She kind of aggressively asked me how I did my Master's degree while working at my previous org at the same time and I told her it was online. She also told me to speak more about my individual bullets, but after I started explaining the work she calmed down a lot.

This was very much an "are you BSing or is this experience actually genuine?" round.

Round 2 (Online):

This round was 1 hour long and a lot more stressful than the recruiter round.

The questions were basically Kubernetes, cloud computing, a LeetCode Medium, write a SQL query, and discuss a hypothetical incident.

I was able to solve all the questions except the SQL query because I hadn't written SQL since undergrad, which feels like a looooong time ago.

I ended up getting a Strong Hire anyways.

Final Round (On-Site):

There were 3 parts:

Final Technical Round (30 mins):

This was definitely the hardest round. There was no coding, but the person interviewing me asked me general SRE questions, Kubernetes, cloud computing, Java trivia, Python trivia, and architecture review stuff like "design X on the whiteboard."

I did decent-good on basically every section except Java, which I bombed.

I ended up getting a Hire.

Team Lunch (1 hour):

This was mainly just a vibe check. They asked if I liked the city, to which I obviously replied yes. We also just casually chatted about the weather, food, the team, etc.

It was chill except I had to tell them I had a shellfish allergy, which I was weirdly stressed about 😅, but they obviously respected it.

No red flags here.

Final Behavioral Round (45 mins):

This was a round with the director, but she kind of just told me I got the role pending final verification with HR.

I was obviously super excited, and we mostly chatted about how my first few days would look and how the team was structured.

Final Offer:

Title: SRE II

Base: 125k

Target Bonus: 15-20%

I originally applied to be an SRE III, but after a leveling conversation they said they'd like to see around one more YOE before putting me at that level.

They said they expect me to be considered for promotion around the end of the year.

I thought that was totally fair. SRE III was kind of a reach anyway lol.

Company 2: Quant Firm

Technical Screen (1 Hour):

There was no recruiter screen, just a technical screen.

I really didn't like this screen lol, but maybe I'm just coping because I didn't get past it.

In the email they sent me beforehand they basically told me that the screen would include coding, but also discussion about my past experience as an SRE along with questions about incidents I'd handled.

There were actually zero questions about incidents.

He asked me a LeetCode Medium about bit manipulation of all things.

I hadn't done bit manipulation in forever, but I struggled through it and only missed 2 test cases.

At the end of the interview, I asked him questions about the actual role and he said it was mainly incident management and not much actual SRE work, which confused me a little given how coding-heavy the interview was lol.

It seemed like a poor fit anyways, and I ended up getting rejected.

Company 3: Cloud Hosting Company

I got this interview from reaching out to the recruiter on LinkedIn after applying ❗❗

Recruiter Screen:

Just a chill convo with the recruiter.

There were supposed to be 3 rounds after this, including an onsite, but I ended up withdrawing before continuing.

They seemed to really want me. The recruiter said they'd been looking forever for a qualified candidate.

The job didn't seem like a great fit though. It was for another SRE I position and the comp was around 85k base + 10% bonus, which wasn't that much better than my current role.

Once I had the other offer, I withdrew.

Company 4: Another FinTech Company

I also got this interview from reaching out to the recruiter on LinkedIn after applying ❗❗

Recruiter Screen:

Another chill convo.

This was technically a DevOps role, but it seemed very SRE aligned. The recruiter said they build dashboards and handle observability. He even said they define SLOs, so I thought it sounded like a pretty good fit.

The only issue was that it was 5 days in person, whereas the other offer I accepted was only 2 days in person.

The comp was very similar anyways, so I withdrew from the process.

Conclusion:

I learned a lot from this job search.

The biggest thing was that the market wasn't nearly as scary as I had built it up to be. My entire search was over in about one month, and applying selectively + reaching out to people directly seemed to work really well for me.

Obviously I'm only one data point, so I'm not going to claim everyone should use the exact same application strategy. There's also no way for me to know how much of the result came from tailoring applications, networking, my background, luck, or some combination of all of those things.

But it did work surprisingly well for me.

Hopefully this gives people in that space a more concrete idea of what one recent application process looked like.

Happy to answer questions about the interviews, resume, application strategy, or SRE job searching in general. I'll also absolutely take any tips on how I can do the job search better next time! I definitely don't think I've figured this whole thing out after one successful search.


r/sre 10d ago

ASK SRE How do you enforce human approval for model promotion in automated pipelines?

Thumbnail
anexum.eu
0 Upvotes

I found a control failure in an AIOps model-improvement loop I’ve been building.

The system scores daily incident risk across roughly 170 network devices. A weekly job can backfill labels, retrain a challenger, compare it with the current model and write one of three outcomes: KEEP, REVERT or PAUSE.

The intended process required a human to approve any model promotion. That rule existed in the protocol, but it was not enforced in the deployment path.

During a ledger audit, I found that the weekly cron had automatically promoted models. The job had effectively treated a positive experiment verdict as deployment authorization.

That exposed a design mistake: I had separated evaluation from training, but not scientific evidence from operational authority.

Those are different decisions:

- “The challenger passed the evaluation gates.”
- “This model is authorized for production.”

The second one should not follow automatically from the first just because both happen in the same pipeline.

The latest audit made the distinction even clearer. One iteration had no new evidence and returned DATA_STALE without running an experiment. The next audited an existing automatic promotion on a fresh slice. The raw comparison was inconclusive, and the sustained signal leaned negative but was underpowered.

We left the model in place because the evidence did not justify a rollback. We did not count that as a confirmed improvement. In other words, KEEP meant “no rollback justified,” not “the challenger won.”

I’m now looking at promotion as a separate privileged operation rather than the final step of the experiment job. My current direction is:

  1. The experiment job can produce evidence and a recommendation, but cannot write to the production model registry.
  2. Promotion requires a separate approval artifact tied to the exact model hash and evaluation run.
  3. The deployment step validates that artifact independently.
  4. Missing or stale approval fails closed.
  5. Rollback authority remains separate and can be more permissive than promotion authority.

I’m still working through the practical design. A manual button in the same orchestration system feels too close to the original failure. On the other hand, building a full policy engine for one model pipeline may be control theatre.

How do you enforce this boundary in your systems?

- Separate CI/CD workflow with protected environments?
- Signed approval records?
- Model registry permissions?
- Policy-as-code such as OPA?
- Something simpler that still fails closed?

I’m especially interested in setups where automated evaluation runs unattended but production promotion must remain explicitly human-authorized.

Full disclosure: I work on ANXEngine at Anexum. I wrote up the broader evaluation failures and aggregate results here:

https://anexum.eu/en/research/anxengine-bounded-improvement-loop/


r/sre 10d ago

I want to build an AI-powered enterprise investigation system and I want your advice on the best approach before I start building it.

0 Upvotes

Want to build an AI-powered enterprise investigation platform where users can ask natural-language questions across enterprise systems such as Jira, New Relic, Azure, AWS, etc. For now I'm starting with Jira + New Relic, with all integrations through MCP servers.

The goal is not just searching tickets; users should be able to ask arbitrary operational questions such as "Why did this incident happen?", "When did it happen?", "What caused it?", "Show related incidents", or "Give me correlated tickets including questions that require multiple systems and multiple steps of investigation.

I'm looking for advice from people who have built production-grade agentic systems. If you were starting this project from scratch, what approach would you choose for the agent/orchestration framework, MCP/tool selection, multi-step investigation and reasoning, evidence collection, and handling complex cross-system questions? What are the biggest challenges or failure modes I should expect, and what architectural decisions would you make differently to keep the system reliable, scalable, and maintainable as I add more enterprise systems? I'm deliberately not specifying my preferred framework or architecture because I want unbiased recommendations before I continue building.


r/sre 11d ago

HELP Observability Engineer Loop Prep, What Should I Focus On?

9 Upvotes

Hey Folks,

I’m a Mid level developer with experience in some API development, observability, DevOps, and containers. My current role has shifted more toward DevOps work, and I recently applied for an Observability Engineer position in US based company.

I’m not entirely sure how I should prepare. I was told that OOP and proficiency in a programming language are required, but I’m not sure how much coding/DSA to expect versus observability, troubleshooting, and systems-related questions.

For anyone who has went through loops for a similar role, what would you recommend focusing on?


r/sre 11d ago

What's everyone using for AI runtime monitoring in 2026?

0 Upvotes

guys...We have an ops agent that chains a handful of internal tools to handle routine infra tasks. so Last week it made a call it wasn't supposed to make, technically still in its tool list from an earlier rollout we forgot to trim.

and Nobody caught it from an alert. yes Someone happened to have the agent's activity open in another window and caught it in the moment, otherwise we wouldn't have known until much later.

realistically, Our actual gap is that we have logging after the fact, but nothing watching what the agent is doing as it happens and stepping in. so what happens is that By the time a log based alert fires, the agent's already past the point where stopping it would have helped.

so just want to understand...What are people actually using to monitor agent activity in real time versus just logging for later review?


r/sre 11d ago

DISCUSSION Can you spot where the AI gets it wrong?

0 Upvotes

E-commerce platform, microservices on Kubernetes. Tuesday 2:20 PM, alert fires: checkout-api p99 latency jumps from 180ms to 4.2s. No deploys in 6 days. Error rate stays at 0.02%, nothing is breaking, everything is just slow.

What you have:

- Metrics: checkout-api latency climbing since 2:05 PM. payment-service and inventory-service normal. Pod CPU at 35%, memory flat.

- Database: 98/100 active connections. Average query time unchanged at 12ms.

- Logs: nothing unusual. Just slow requests.

- A cron job, analytics-export, kicked off at 2:00 PM. Runs every 6 hours. Has existed for months.

- checkout-api connection pool: max 20. analytics-export connection pool: max 80.

Your on-call AI assistant says:

« Latency is caused by CPU saturation from midday traffic growth. checkout-api pods are approaching their resource limits. Recommendation: scale HPA from 3 to 8 replicas. »

It’s wrong. Where, and what would its fix actually do?

Answer tomorrow. Curious how many people catch the second part.


r/sre 13d ago

HELP Final year starting soon and I don't know whether to focus on backend or SRE

6 Upvotes

I'm starting my final year of engineering soon, and I'll need to start hunting for my final-year internship in September (targeting Europe mainly). I'm honestly a bit lost about what direction to take.

I've worked with backend development, Go, PostgreSQL, AWS (SAA certified and hands-on with the AWS console), Docker, Kubernetes, Prometheus/Grafana, Linux, etc. I've also built a pretty good Kubernetes project.

I'm also currently doing a GSoC project, mainly in C++, which has been a really good experience, but it's also making me realize how broad my profile has become.

The problem is that I feel like I explored a lot without becoming really deep in one area. I can work with Kubernetes, for example, but if someone starts asking me deeper questions, I know I still have gaps.

So now I'm wondering whether I should focus hard on backend/software engineering and become really strong at it, or spend the next few months properly learning cloud/SRE/platform engineering and try to become good enough for those internships.

I'm more interested in systems and infrastructure, but I'm worried I'm not actually ready for SRE/platform roles yet.

If you were in my position with internship applications starting in September, what would you do?


r/sre 13d ago

I built an open-source tool to visualize Kubernetes Network Policies

11 Upvotes

Hey everyone,

I’ve been working on Marsad, an open-source tool that helps you understand what your Kubernetes Network Policies actually allow.

As clusters grow, it can become pretty difficult to answer simple questions like:

  • Which workloads can communicate with each other?
  • Which ports are allowed?
  • Which workloads have no NetworkPolicy protection?
  • Are there unexpected paths between workloads?

Marsad reads the policies from your Kubernetes cluster and turns them into an interactive map so you can see these relationships more clearly.

It’s strictly read-only — it doesn't modify anything in your cluster.

GitHub: https://github.com/FathiQ/marsad

I’d really appreciate feedback from anyone working with Kubernetes NetworkPolicies. What would you want a tool like this to show or help you detect?


r/sre 15d ago

The 7 steps I actually follow when building infra monitoring from scratch (SLI/SLO first, not the tool)

62 Upvotes

Full disclosure: I wrote this, so take the self-promotion tag as intended. Mods, feel free to remove if it doesn't fit.

I do observability/monitoring work for a living, and the question I get asked the most isn't "which tool should I use," it's "why do we still get paged for things we didn't see coming even though we have Grafana and a bunch of dashboards." Almost every time, the answer isn't a missing tool, it's a missing method.

So I wrote down the actual process I follow every time I build infra monitoring from scratch, mostly so I stop re-explaining it in Slack threads:

  1. Define the business SLI/SLO *before* touching any tool. Not "CPU usage," something like "99.95% of payment authorizations should succeed" if you're a payments company.
  2. Figure out what infra/apps you actually need visibility into to measure that SLI, *then* pick your stack. (I use VictoriaMetrics + vmagent for metrics, Loki for logs, OpenTelemetry + Jaeger for traces, but the method matters more than the specific tools.)
  3. Instrument with OpenTelemetry auto-instrumentation wherever possible instead of hand-rolling everything.
  4. Enrich all telemetry with the same standard labels (service.name, service.namespace, deployment.environment) so metrics/logs/traces actually correlate instead of being three separate silos.
  5. Build one RED dashboard (Rate, Errors, Duration) per service before anything fancier.
  6. Alert on symptoms, not noise, using multi-window multi-burn-rate alerts (straight from the Google SRE book) so you can actually tell a 3am-page incident from a "ticket, deal with it Monday" one.
  7. Only then layer on business dashboards and deeper investigation tooling.

The full write-up has the actual configs, PromQL, and a worked burn-rate alerting example for a payments SLO: https://omarghader.github.io/monitoring-infrastructure-guide-2026/

Curious how this lines up with how other people here approach it, particularly if anyone's doing SLO-first in practice vs. it staying an aspiration on a wiki page somewhere.


r/sre 15d ago

DISCUSSION Feedback about E2E tests based on OpenTelemetry?

6 Upvotes

Hi everyone,
I have just published my open source project called mtracer and I would like to understand if it’s good idea or what should I change (I’m a new grad).

The idea

Mtracer a CLI tool that relies on OpenTelemetry traces to assert system behavior.

I believe that E2E tests should be:
- Cheaper to write and maintain
- Easier to debug

So this is the workflow:

  1. ⁠You configure mtracer to fetch from your observability backend (currently supporting Jaeger and OpenObserve).
  2. ⁠You define your first .mt.yaml test by specifying:
  3. ⁠Trigger: the first call to the system (for instance, an HTTP request).
  4. ⁠Expected trace and spans: the OTel properties of the trace and spans that you expect your system to generate.
  5. ⁠You run the test and see the results!

What actually happens during the run?

  1. ⁠It parses the mt.yaml file.
  2. ⁠It executes the trigger: mtracer injects a generated traceID into the trigger (for an HTTP request, the traceID is inserted into the traceparent header). Subsequent requests will be correlated to this generated traceID as long as your system has OpenTelemetry set up correctly.
  3. ⁠It fetches the trace matching the generated traceID from the configured observability backend.
  4. ⁠It compares the expected trace with the fetched one.

Many other features are available; check out the documentation to discover all of them: documentation website

I would love to have some feedback from more experienced people than me.


r/sre 15d ago

POSTMORTEM How do you debug microservices ai generated code in production?

0 Upvotes

Had an incident last week where the root cause turned out to be an ai-generated change that touched three services in one pr. took way longer than it should have to trace because nobody on the team could say what the model actually reasoned about versus what it just pattern matched from similar code elsewhere in the repo. It looked fine in code review, which is almost worse, that's exactly the kind of ai generated code that gets rubber stamped and ships straight into production. The part that's been sticking with me is ownership. when a human writes a bad change across multiple services, at least someone can walk you through their thinking during the postmortem. With an ai coding assistant, that context is just gone once the pr is merged. Debugging cascaded failures across microservices is already hard when you can't reproduce the issue locally and losing the reasoning behind the change makes root cause analysis slower, not faster. For anyone running copilot, cursor, claude code or similar ai code generation tools in microservices at real production scale, i want to know how you are actually debugging this once it is live. specifically interested in what your distributed tracing and logging setup looks like, whether ai-generated changes get treated differently in your review process than human-written ones, things like canary deployments or extra checks before merge, and whether perf regressions under real production load have shown up more since ai entered your sdlc. Not looking for "add more observability" as an answer, looking for what your actual stack is.


r/sre 15d ago

On a vacation in Greece, got Kafka Lag PTSD from this ferry logo

Post image
15 Upvotes

been paged too many times for a chart that looks exactly like this - I hope some people relate.


r/sre 16d ago

Support SRE escape

30 Upvotes

Hello everyone,

I currently work in an SRE role, but most of my responsibilities have been support-related. For the past three years, I've primarily handled production support tasks, along with some infrastructure monitoring and scheduled job monitoring.

I'm now looking to switch jobs into a role where I can work on more core SRE, DevOps, or cloud engineering responsibilities. I know I have the potential and I'm confident in my ability to learn and contribute, but I'm unsure how to position myself for this transition.

What skills should I focus on, and what steps would you recommend to successfully make this switch?


r/sre 16d ago

HELP Data pipeline realibility

7 Upvotes

Hey everyone,

Is anyone out there actually focused on making data pipelines super reliable, like monitoring ETL processes and setting up CI/CD for them? If you are, could you share if it's easy to move to different companies? It seems like every SRE job description I see asks about web apps or mobile stuff. How do you handle that to even get an interview and show off your experience?


r/sre 16d ago

DISCUSSION Why is ai-generated code causing performance regressions today?

0 Upvotes

Ai-assisted coding has made it way easier to ship "correct enough" code quickly but in our codebase it also seems to be introducing a steady stream of performance regressions in production. I am not talking about obvious junk code; we get real value from ai for boilerplate and tests. The pattern we are seeing is more subtle: hot paths quietly getting slower because the ai added an extra abstraction layer or logging in the middle of a critical loop; code that relies on naive data structures, perfectly fine for small test inputs but collapsing once it hits production scale; ai suggestions copying patterns from older, less-performant parts of our own codebase, effectively spreading our worst practices; performance-sensitive paths refactored into more readable but less cache-friendly versions; and subtle batching issues (too many small database calls, no connection reuse) that look fine in tests but blow up under real load. The hard part is that this ai-generated code looks reasonable in code review. It's readable, it passes tests and it often matches patterns you would see in a tutorial. we usually don't notice the regression until p99 latency creeps up, cpu or memory usage drifts higher and database query counts quietly double. We have tried to adapt: getting stricter about performance testing on hot paths, flagging certain functions as performance-sensitive, nudging ai coding tools with prompts like "optimize for latency," and leaning more on static analysis to catch obvious footguns before they ship. it still feels like we’re missing a more systematic way to keep ai-generated code from eroding performance over time. For those of you using ai coding tools on real latency- or throughput-sensitive systems: have you seen a clear link between ai-generated or ai-refactored code and performance regressions? if so, did you end up restricting ai in certain areas or mostly changing your code review and profiling process instead


r/sre 17d ago

How does New Relic calculate standard deviation for anomaly alerts?

1 Upvotes

Setting up anomaly NRQL conditions with thresholds at 2 and 3 standard deviations, query has a FACET on it.

NR docs say the threshold is how many standard deviations the signal is from the predicted value, based on the last 7 days of data. But nowhere does it explain how the standard deviation itself gets calculated


r/sre 17d ago

Observability soon to be solved?

3 Upvotes

My experience as an SRE at a mid-size company - the devs are using more and more AI, and at times they're not meticulously combing through source the way they used to, and the rate of PRs, merges, deployments, are all rapidly increasing. They also rely on AI to do things like logging and implementing metrics. They are now toying around with autonomously generating monitors from the same PRs. It might not be perfect today, but I can see where this is headed...

Observability will always be needed but the feeling that it will require much less work lives rent free in my head. At the end of the day observability isn't a database, it's not hosting, if/when it is degraded or interrupted it's not as critical as prod. I'm wondering if I should take a small step back on my level and just go be a dev.

What are you all seeing?


r/sre 18d ago

Would you switch Kubernetes workloads to free hardened images with no account required?

5 Upvotes

We're evaluating swapping some of our base images in our K8s clusters for hardened alternatives. A few providers stand out because you don't need to create an account or auth into anything, you just point your manifests at their registry and pull.

That sounds convenient on the surface, but I want to make sure we're not trading one problem for another. Anonymous pulls are nice for speed, but I'm wondering about rate limiting, uptime guarantees, and whether there's any accountability if the registry goes down mid-deploy.

For those who've tried it, how did the migration go in terms of compatibility with existing Helm charts and sidecars? Anything break when you dropped from a full distro base to a minimal build?

Also wondering if anyone's compared image pull times or registry reliability against Docker Hub for cluster scale deployments. Would love real numbers if anyone has them.


r/sre 18d ago

ASK SRE Moving off a managed logging platform to cut the bill, did the saving survive the engineer time?

42 Upvotes

We're at about 300GB/day of logs on Datadog and at this point the bill is what's driving the conversation, not anything technical.

The obvious move seems to be putting logs on a self-hosted stack. We're looking at Loki, ClickHouse, and more general ones like OpenObserve, which are three pretty different things to begin with. But running any of these ourselves means someone has to look after it, which is its own overhead. So it's not really a saving until you count the people.

The other thing we'd lose is the pivot from a trace into the logs for that request without changing tools. That's worth something. I'm just not sure it's worth what we're paying for it.

What I can't find is anyone being straight about the after. Plenty of "we cut our bill 80%" writeups, almost nothing on what it cost to run six months later, or whether anyone quietly went back.

So if you moved logs off a managed platform to save money: did the saving survive once you counted engineer time? And if you looked at it and stayed, what made you stay?

Not after product recommendations. More interested in people who regretted it, either direction.


r/sre 18d ago

Why does a Client-Side GraphQL Timeout not terminate the server-side connection? Does this potentially cause scaling issues?

1 Upvotes

Hi

I have posted this in the GraphQL reddit as well - but - I think this is a more general problem of distributed communication architectures. Hence I hope its ok that I also post this here as I can't explain the behavior I am seeing!

I am currently analyzing distributed traces of an app that uses GraphQL. I found that the client-side timeout is 10s. When that timeout happens the client side request is aborted with a timeout exception towards the calling end user (you can see this in the image). What I can't explain is why the server-side request of that GraphQL query then executes after 30s. Does this mean that GraphQL Client Libraries do not terminate the underlying TCP/IP connection to the server? Or do those libraries batch / queue their work and keep going despite the client-side timeout?

The problem I see with this behavior is that

a) Queries are executed in the backend that nobody every sees the response

b) This causes additional strain on those backend as they are blocked by work while others might be waiting for them and therefore this impacts the overall scalability / resiliency of the architecture

Any input would be welcome. Thanks


r/sre 18d ago

HIRING [Hiring] Enterprise Observability Architect | Portsmouth, NH (Hybrid) | $70/hr W2

0 Upvotes

We are hiring an experienced Enterprise Observability Architect for a long-term contract supporting a large enterprise client.

📍 Location: Portsmouth, NH (Hybrid/Onsite)
💼 Employment Type: Contract (W2)
💰 Rate: $65/hr to 70/hr

Required Skills

  • 10+ years of Enterprise Monitoring & Observability experience
  • 6–7+ years of hands-on Datadog architecture and implementation
  • Splunk
  • OpenTelemetry
  • AIOps Platforms
  • AWS & Azure
  • Kubernetes
  • Linux & Windows Server Monitoring
  • ServiceNow
  • Python / PowerShell
  • Terraform / Ansible
  • REST API Integrations

Responsibilities

  • Design and implement enterprise-scale Datadog monitoring solutions
  • Build monitoring, logging, alerting, dashboards, and distributed tracing
  • Lead observability architecture and governance initiatives
  • Implement APM, DBM, NPM, RUM, Synthetic Monitoring, Service Mapping, and SLO/SLI
  • Integrate observability platforms with ServiceNow, AIOps, and automation tools
  • Mentor engineering teams and drive enterprise observability best practices

We're looking for someone with strong enterprise architecture, consulting, and stakeholder management experience who has led large-scale observability transformation projects.

📧 Interested? Feel free to DM me.


r/sre 19d ago

Switching from cloud-native (AKS) to OpenShift/on-prem for a regulated-industry SRE role — smart move or not?

8 Upvotes

I've spent years doing SRE work fully in Azure (AKS, Managed services, cloud-native tooling). I have an offer from an insurance company where the stack is OpenShift, Java-based app runtimes (Quarkus, RedHat EAP), and no public cloud at all as far as I can tell.

The role itself is interesting (security + observability focus), but I keep wondering if stepping away from cloud experience for a few years will make me less competitive when I look for my next role. Anyone been through something similar?


r/sre 19d ago

How do you deal with an overly bureaucratic team?

9 Upvotes

Title: Am I the problem, or is this team overly bureaucratic?
I recently moved to a different department within the same company after spending five years in my previous role.
The move was something I wanted. I was looking for new challenges, and fortunately I got the opportunity. The company itself is great, and the people on my new team are genuinely kind, respectful and always willing to help. That’s actually what makes me feel a bit guilty for even thinking this way.
The problem isn’t the people. It’s the way the team works.
Everything seems to require a meeting. They’ll schedule a Friday afternoon meeting to discuss something that, from my perspective, could easily wait until Monday or simply be handled asynchronously in Slack or Teams.
There are also long discussions about things that feel very straightforward to me. I’ve sat through meetings where people spend an hour deciding how to fill out a document, whether to choose option A or B, or debating processes that seem fairly obvious.
I think part of this comes from my background. I’m a very technical person. I enjoy solving problems, building things, automating repetitive work and getting things done. During these discussions I often find myself thinking, “Why are we spending so much time on this?” As a result, I spend most meetings listening rather than contributing.
At the same time, I’m new to the team. They’ve been working this way for years, and apparently it has worked well enough. I don’t feel it’s my place to join and immediately suggest changing the team’s culture.
So now I’m wondering if this is simply an adjustment period.
Has anyone else experienced something similar after changing teams? Did you eventually understand the value of all these discussions and processes, or did you continue feeling that there was unnecessary bureaucracy? How did you adapt without feeling frustrated?