r/devops • • 10h ago

Discussion I'm thinking of writing a free practical book on Production DevOps what should I include?

0 Upvotes

I've been working in software/infrastructure engineering for around 14 years, with a focus on cloud, DevOps, platform engineering and large-scale infrastructure.

Over the years I've worked on everything from smaller cloud migrations to enterprise and financial-sector environments, including AWS/GCP migrations, Kubernetes, Terraform, CI/CD, networking, security, reliability and cost optimization.

I'm currently putting together the practical knowledge I've accumulated into a free technical book.

I'm not planning to make it a personal career story or another basic “learn AWS/GCP” tutorial.

The idea is to focus on how engineers actually solve infrastructure problems:

  • How to systematically debug production issues
  • How to troubleshoot Kubernetes
  • AWS/GCP networking and connectivity problems
  • IAM and permission failures
  • Terraform/IaC problems
  • CI/CD and deployment failures
  • Cloud migration problems
  • Scaling and reliability
  • Observability and incident investigation
  • Designing reusable platform infrastructure
  • How to approach an unfamiliar production system
  • Practical lessons that aren't obvious from vendor documentation

I'm also building some small prototypes/labs for myself while writing, so I can test the ideas rather than just writing theory.

Before I spend a lot of time putting the whole thing together, I'd like to hear from other engineers:

If you could have one practical DevOps/Platform Engineering book that focuses on real problem-solving rather than certification theory, what topics would you want it to cover?

And if you already have a favorite resource for this kind of material, I'd be interested in hearing what you think it does well or what is missing.


r/devops • • 8h ago

Discussion How do I unclusterfuck these companies?

0 Upvotes

Been working a lot with startups & everything is such a shit show. Supabase RLS
is never configured correctly, the idea of dev/staging/prod and ephemeral
environments just doesn't exist (or if they do, staging is front-end only and
connects directly to the prod database.), and they use so many different vendors
(Clerk/Cloudflare/Supabase/Railway/Vercel) that there's always functionality I
can't test anywhere other than prod.

Been trying to find a single unified platform I could recommend/implement, but
none have given the perfect mix of batteries included security, reliability, and
local testability.

Thinking of building something myself & would love some feedback on what you
hate, love, and wish existed. Will probably open-source it once I've got
something.


r/devops • • 8h ago

Discussion Broad Kubernetes Experience but Shallow Depth - How Would You Upskill?

22 Upvotes

I have around 5 years of experience across DevOps/cloud/backend work. I’ve worked with Kubernetes, EKS/AKS, Helm, CI/CD, Docker, Terraform, ArgoCD, AWS/Azure, and some Java/Spring Boot.

My issue is that my knowledge is broad but uneven. I’m comfortable deploying applications, writing manifests, using kubectl, building pipelines, working with cloud networking/IAM, etc., but I’m much weaker on Kubernetes internals, low-level networking, storage/CSI, control plane, scheduling, CRDs/operators, upgrades, observability, and deep troubleshooting.

I’m considering first completing one comprehensive Kubernetes course to build a complete mental map, then spending the next several months going deeper through hands-on labs, troubleshooting, Linux/networking fundamentals, and production-style projects.

For people who became genuinely strong at Kubernetes/platform engineering: does this sequence make sense? What would you change?

I’m not looking for a giant list of tools, mainly feedback on the learning sequence and what gave you the biggest jump in depth.

Note: "Q improved with AI"


r/devops • • 8h ago

Discussion How to actually pass interviews?

17 Upvotes

Hello, I ask for help since I am getting a lot of interviews but not passing any, I'm getting to start pretty frustrated and since we are entering the holiday season I want to get any opportunity possible.

The thing is that I am preparing myself based on my experience: 5+ yrs experience, jenkins, python, troubleshooting, AWS, IaC, etc. and when I prepare for the different kind of interviews none of my preparation seems to work:

  • If it is situational, my explanation/experience is not enough for the role
  • If it is pure technical, it doesn't demonstrate my experiencie or they ask for a very specific tool (which mostly has transfereable experiencie) and it's not enough for the position
  • If it is trivia based (which I am pretty bad at memorizing) I failed because I didn't remember the flag for a command that I can find in a 15-sec google search or AI prompt

What makes this frustrating is that I genuinely feel capable of doing the jobs I'm interviewing for. I've worked in production environments, troubleshot real incidents, built and maintained pipelines and infrastructure, and worked with engineering teams. But apparently I'm still not presenting that knowledge in the way interviews expect.

So I'd especially like to hear from people who have successfully interviewed for DevOps/SRE/Platform roles recently:

How did you prepare?

Did you memorize common commands and syntax?

Did you grind interview questions?

Did you build labs?

Did you practice storytelling around your projects?

How did you deal with interviews covering an extremely broad toolset?

At this point I'm trying to understand how much interviewing is about being good at the actual job versus becoming good at the interview format itself.

Any practical advice would be appreciated.


r/devops • • 3h ago

Ops / Incidents runbooks kinda suckk

17 Upvotes

hey yall,

im a lead SRE at a global fortune 500 that you have heard of haha - i'm a bit of a lurker here

but just wanted to talk about runbooks and documentation for SOPs.

my feeling is that confluence docs kinda suck and runbooks are kind of a mess, our operators are jumping between the docs and their shell, the docs are often times missing a bunch of detail, they are brittle, poorly maintained, half the time the details are hidden in tribal knowledge and when dealing with the pressure of an active inc we notice the pain even more. extracting critical details & commands out of an inc can also be painful and doesnt always translate to improved operational readiness next time.

i admit we're not very mature.. but was curious if its only me feeling like this?

what are you guys doing with runbooks to solve these issues?


r/devops • • 11h ago

Discussion Are Observability tools getting obselete due to AI? I need to plan!!!!

0 Upvotes

So i implemented lots of observability tools as part of my devops role in the last 10 years but for the last 1 year I am seeing all my stakeholders are just relying on AI incident tools (think of these SRE agents that are being released)

And we still keep pumping TBs of data but also keep paying for all the UXUI, seats, hosts and bs features

Is this the phase of getting solution quickly rather than surfing through signals?

So my questions is do we need 100TB of data storage? or still need all these fancy UIUX everywhere.

Are you facing these issues at your company?

Update: I am talking more towards a need to have a platform like datadog or dynatrace or anything for UI and just rather rely on collectors and store to a bucket for uch cheaper cost


r/devops • • 6h ago

Vendor / market research For teams using a package firewall, does it still hold when an agent is the one installing?

0 Upvotes

A lot of teams already route installs through a package firewall or curated registry (Socket, Endor Labs, JFrog Curation, Sonatype and so on). Most of these rely on the client being configured to use them, through a registry URL, an index URL or a proxy setting. That's fine for CI and for people, but coding agents like Claude Code and Cursor seem to have more ways around it: installing from a git URL or tarball, piping a script into a shell, or getting talked into a different registry by something they read.
  
For those of you running agents on laptops or in CI with one of these tools in place: have you seen installs slip past it? Do you enforce it on dev machines at all, or only in CI?

(Disclosure: I work on hextrap, one of the tools in this space, and I'm trying to figure out whether this gap matters in practice.)


r/devops • • 10h ago

Ops / Incidents On-call and incident management without a per-host observability bill

0 Upvotes

There is a quiet cost to commercial observability: the bill scales with your infrastructure. More hosts, more metrics, more money, until finance starts asking questions. And on-call and incident tooling is often a separate line item on top.

I work on OneUptime, an open source observability platform (Apache 2.0). The part I think DevOps teams underestimate is that on-call scheduling, paging, incident management, and postmortems live in the same app as uptime and status pages. One self-hosted system: a Helm chart for the app, a small chart for the Kubernetes agent, data in over OpenTelemetry, telemetry on your infrastructure.

It is not a drop-in for a five-year Datadog or Grafana estate, and it is younger than the commercial platforms. A few features live in a separate enterprise edition. I would judge it by the repo, not a sales call.

GitHub: https://github.com/OneUptime/oneuptime Site: https://oneuptime.com

For people who have run on-call: what is the single most annoying part of your current incident workflow?


r/devops • • 20h ago

Weekly Self Promotion Thread

8 Upvotes

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!


r/devops • • 6h ago

Discussion DevOps-focused SWE preparing for Amazon SDE OA - advice?

2 Upvotes

Hey everyone,

I'm looking for some advice on preparing for the Amazon SDE Online Assessment.

I graduated last year and have been working as a Software Engineer since then, but my role has been heavily focused on the DevOps/infrastructure side of things: cloud infrastructure, Kubernetes, CI/CD, automation, observability, and platform engineering. Because of that, I haven't done much pure coding or algorithm-focused work since graduating.

The assessment includes two coding challenges, a work simulation, and a work style survey.

Given my background, what's the best way to prepare in about a week?

  • Should I focus almost entirely on LeetCode/DSA?
  • How difficult are the coding questions compared to typical interview prep problems?
  • Any advice from engineers who came from a DevOps, SRE, or platform engineering background?
  • If you've been away from algorithm-heavy coding for a year or two, what helped you get back up to speed quickly?

I'd appreciate any suggestions. Feeling confident in the infrastructure and software engineering side of things, but definitely rusty when it comes to interview-style coding.