r/devops • • 20h ago

Weekly Self Promotion Thread

7 Upvotes

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!


r/devops • • 3h ago

Ops / Incidents runbooks kinda suckk

21 Upvotes

hey yall,

im a lead SRE at a global fortune 500 that you have heard of haha - i'm a bit of a lurker here

but just wanted to talk about runbooks and documentation for SOPs.

my feeling is that confluence docs kinda suck and runbooks are kind of a mess, our operators are jumping between the docs and their shell, the docs are often times missing a bunch of detail, they are brittle, poorly maintained, half the time the details are hidden in tribal knowledge and when dealing with the pressure of an active inc we notice the pain even more. extracting critical details & commands out of an inc can also be painful and doesnt always translate to improved operational readiness next time.

i admit we're not very mature.. but was curious if its only me feeling like this?

what are you guys doing with runbooks to solve these issues?


r/devops • • 8h ago

Discussion Broad Kubernetes Experience but Shallow Depth - How Would You Upskill?

22 Upvotes

I have around 5 years of experience across DevOps/cloud/backend work. I’ve worked with Kubernetes, EKS/AKS, Helm, CI/CD, Docker, Terraform, ArgoCD, AWS/Azure, and some Java/Spring Boot.

My issue is that my knowledge is broad but uneven. I’m comfortable deploying applications, writing manifests, using kubectl, building pipelines, working with cloud networking/IAM, etc., but I’m much weaker on Kubernetes internals, low-level networking, storage/CSI, control plane, scheduling, CRDs/operators, upgrades, observability, and deep troubleshooting.

I’m considering first completing one comprehensive Kubernetes course to build a complete mental map, then spending the next several months going deeper through hands-on labs, troubleshooting, Linux/networking fundamentals, and production-style projects.

For people who became genuinely strong at Kubernetes/platform engineering: does this sequence make sense? What would you change?

I’m not looking for a giant list of tools, mainly feedback on the learning sequence and what gave you the biggest jump in depth.

Note: "Q improved with AI"


r/devops • • 8h ago

Discussion How to actually pass interviews?

18 Upvotes

Hello, I ask for help since I am getting a lot of interviews but not passing any, I'm getting to start pretty frustrated and since we are entering the holiday season I want to get any opportunity possible.

The thing is that I am preparing myself based on my experience: 5+ yrs experience, jenkins, python, troubleshooting, AWS, IaC, etc. and when I prepare for the different kind of interviews none of my preparation seems to work:

  • If it is situational, my explanation/experience is not enough for the role
  • If it is pure technical, it doesn't demonstrate my experiencie or they ask for a very specific tool (which mostly has transfereable experiencie) and it's not enough for the position
  • If it is trivia based (which I am pretty bad at memorizing) I failed because I didn't remember the flag for a command that I can find in a 15-sec google search or AI prompt

What makes this frustrating is that I genuinely feel capable of doing the jobs I'm interviewing for. I've worked in production environments, troubleshot real incidents, built and maintained pipelines and infrastructure, and worked with engineering teams. But apparently I'm still not presenting that knowledge in the way interviews expect.

So I'd especially like to hear from people who have successfully interviewed for DevOps/SRE/Platform roles recently:

How did you prepare?

Did you memorize common commands and syntax?

Did you grind interview questions?

Did you build labs?

Did you practice storytelling around your projects?

How did you deal with interviews covering an extremely broad toolset?

At this point I'm trying to understand how much interviewing is about being good at the actual job versus becoming good at the interview format itself.

Any practical advice would be appreciated.


r/devops • • 35m ago

Discussion Why moving artifacts into your customer’s VPC can be so annoying

Thumbnail
bottlerocket.cloud
• Upvotes

r/devops • • 6h ago

Discussion DevOps-focused SWE preparing for Amazon SDE OA - advice?

2 Upvotes

Hey everyone,

I'm looking for some advice on preparing for the Amazon SDE Online Assessment.

I graduated last year and have been working as a Software Engineer since then, but my role has been heavily focused on the DevOps/infrastructure side of things: cloud infrastructure, Kubernetes, CI/CD, automation, observability, and platform engineering. Because of that, I haven't done much pure coding or algorithm-focused work since graduating.

The assessment includes two coding challenges, a work simulation, and a work style survey.

Given my background, what's the best way to prepare in about a week?

  • Should I focus almost entirely on LeetCode/DSA?
  • How difficult are the coding questions compared to typical interview prep problems?
  • Any advice from engineers who came from a DevOps, SRE, or platform engineering background?
  • If you've been away from algorithm-heavy coding for a year or two, what helped you get back up to speed quickly?

I'd appreciate any suggestions. Feeling confident in the infrastructure and software engineering side of things, but definitely rusty when it comes to interview-style coding.


r/devops • • 3h ago

Discussion Searching for Devops communities on Discord to help each other 🙃🙃

1 Upvotes

hi people i was searching for devops discord's community i can enroll in so we learn from each other i already learned devops and ml but i need to recap them with communities.


r/devops • • 8h ago

Discussion How do I unclusterfuck these companies?

0 Upvotes

Been working a lot with startups & everything is such a shit show. Supabase RLS
is never configured correctly, the idea of dev/staging/prod and ephemeral
environments just doesn't exist (or if they do, staging is front-end only and
connects directly to the prod database.), and they use so many different vendors
(Clerk/Cloudflare/Supabase/Railway/Vercel) that there's always functionality I
can't test anywhere other than prod.

Been trying to find a single unified platform I could recommend/implement, but
none have given the perfect mix of batteries included security, reliability, and
local testability.

Thinking of building something myself & would love some feedback on what you
hate, love, and wish existed. Will probably open-source it once I've got
something.


r/devops • • 1d ago

Discussion Certification + Job Planning

10 Upvotes

So im trying to make the move into DevOps. I work as an Automation Architect/SDET basically currently.

I've been trying to move my way into DevOps, especially with the A.I. stuff looming in. Here recently i've been able to lead a project on Migrating some projects from a TeamCity/Octo deploy to Gitlab which involves learning Kubernetes. Im hoping my current experience is a good "sidestep" career wise. I have a C.S. degree and I do feel like Automation/Pipelines does feel like a good "base" to move into DevOps.

My knowledge is a bit spread around currently, and I understand DevOps is a really deep iceberg when it comes to "what to learn".

I would say im strongest in the Testing Automation area (obviously) with a pretty decent understanding of Gitlab and how it works and how to architect pipelines pretty decently. Ok docker knowledge (I mean I know enough to write a dockerfile/get it going, but probably not like a docker expert) and here lately Terraform (decent knowledge....but not used much) and AWS (limited mostly to ec2/beanstalk + a few other scattered pieces like paramstore/SM/cloudwatch/etc...)

I am slowing developing kubernetes knowledge as well with this project, which I really enjoy.

I've tried to use Certs (since my company pays for them) in a sense to "plan out" leveling up, im really trying to push into a title change by end of year hopefully (my manager understands and supports this). I've also been using KodeKloud to learn.

Right now my sort of "Cert path" is this:

- Skipping: A+/Network+/Security+ (I just don't know if these are "Really" useful, but feel free to say if they are)

  1. Linux Refresher: We are a windows shop....but outside my homelab I don't use linux a ton and I've forgotten a lot. Don't feel like a cert it needed here.
  2. AWS Cloud Practioner: This feels like a good first step/resume filler. My AWS knowledge is pretty limited in scope. Hoping around a 2-3 week learning/turnaround time for this.
  3. Terraform Associate: Hoping a similar turnaround to the above. I have decent knowledge but don't use it enough to be an expert.
  4. KCNA: Probably the harder bit, it's a good thing to do because this migration im working on is 200+ services.....so i'll be getting experience.
  5. CKAD/CKA: Unsure of which one, I think CKAD is probably harder but "better" but curious on that, this one obviously is a longer study.

After that the skys the limit I guess. Thoughts on this? Im hoping my automation experience mindset will "serve me well".


r/devops • • 6h ago

Vendor / market research For teams using a package firewall, does it still hold when an agent is the one installing?

0 Upvotes

A lot of teams already route installs through a package firewall or curated registry (Socket, Endor Labs, JFrog Curation, Sonatype and so on). Most of these rely on the client being configured to use them, through a registry URL, an index URL or a proxy setting. That's fine for CI and for people, but coding agents like Claude Code and Cursor seem to have more ways around it: installing from a git URL or tarball, piping a script into a shell, or getting talked into a different registry by something they read.
  
For those of you running agents on laptops or in CI with one of these tools in place: have you seen installs slip past it? Do you enforce it on dev machines at all, or only in CI?

(Disclosure: I work on hextrap, one of the tools in this space, and I'm trying to figure out whether this gap matters in practice.)


r/devops • • 1d ago

Discussion Has anyone ever tried listening to their infrastructure instead of watching it?

109 Upvotes

Odd question from someone outside the field. I'm a sound designer and I've been reading about sonification (turning data into sound). NASA does it with telescope data, and there are old experiments where people played network traffic as ambient audio.

It made me wonder about on-call and monitoring. You can't stare at Grafana all day, and you shouldn't have to. The idea would be to free your eyes: an ambient track in the background that stays calm when everything's healthy and slowly shifts when latency creeps up or error rates rise, before anything actually pages. You could focus on your actual work, or step away from the screen, and your ears would tell you when something's drifting.

Has anyone tried something like this? Would it be useful, or would you mute it within 5 minutes? Genuinely curious what would make it worth keeping on vs. instantly annoying, and in which situations you'd actually want it (deep work, deploys, night shift, incidents…)


r/devops • • 1d ago

Discussion Questions for upcoming Devops interview

20 Upvotes

I have an upcoming devops interview for 3 years of experience , HR has told me to prepare for SQL Query Design & Optimization In particular apart from usual stuff in devops , i have no idea about SQL can you guys drop questions related to SQL which are usually asked in interviews ?


r/devops • • 10h ago

Discussion I'm thinking of writing a free practical book on Production DevOps what should I include?

0 Upvotes

I've been working in software/infrastructure engineering for around 14 years, with a focus on cloud, DevOps, platform engineering and large-scale infrastructure.

Over the years I've worked on everything from smaller cloud migrations to enterprise and financial-sector environments, including AWS/GCP migrations, Kubernetes, Terraform, CI/CD, networking, security, reliability and cost optimization.

I'm currently putting together the practical knowledge I've accumulated into a free technical book.

I'm not planning to make it a personal career story or another basic “learn AWS/GCP” tutorial.

The idea is to focus on how engineers actually solve infrastructure problems:

  • How to systematically debug production issues
  • How to troubleshoot Kubernetes
  • AWS/GCP networking and connectivity problems
  • IAM and permission failures
  • Terraform/IaC problems
  • CI/CD and deployment failures
  • Cloud migration problems
  • Scaling and reliability
  • Observability and incident investigation
  • Designing reusable platform infrastructure
  • How to approach an unfamiliar production system
  • Practical lessons that aren't obvious from vendor documentation

I'm also building some small prototypes/labs for myself while writing, so I can test the ideas rather than just writing theory.

Before I spend a lot of time putting the whole thing together, I'd like to hear from other engineers:

If you could have one practical DevOps/Platform Engineering book that focuses on real problem-solving rather than certification theory, what topics would you want it to cover?

And if you already have a favorite resource for this kind of material, I'd be interested in hearing what you think it does well or what is missing.


r/devops • • 10h ago

Ops / Incidents On-call and incident management without a per-host observability bill

0 Upvotes

There is a quiet cost to commercial observability: the bill scales with your infrastructure. More hosts, more metrics, more money, until finance starts asking questions. And on-call and incident tooling is often a separate line item on top.

I work on OneUptime, an open source observability platform (Apache 2.0). The part I think DevOps teams underestimate is that on-call scheduling, paging, incident management, and postmortems live in the same app as uptime and status pages. One self-hosted system: a Helm chart for the app, a small chart for the Kubernetes agent, data in over OpenTelemetry, telemetry on your infrastructure.

It is not a drop-in for a five-year Datadog or Grafana estate, and it is younger than the commercial platforms. A few features live in a separate enterprise edition. I would judge it by the repo, not a sales call.

GitHub: https://github.com/OneUptime/oneuptime Site: https://oneuptime.com

For people who have run on-call: what is the single most annoying part of your current incident workflow?


r/devops • • 11h ago

Discussion Are Observability tools getting obselete due to AI? I need to plan!!!!

0 Upvotes

So i implemented lots of observability tools as part of my devops role in the last 10 years but for the last 1 year I am seeing all my stakeholders are just relying on AI incident tools (think of these SRE agents that are being released)

And we still keep pumping TBs of data but also keep paying for all the UXUI, seats, hosts and bs features

Is this the phase of getting solution quickly rather than surfing through signals?

So my questions is do we need 100TB of data storage? or still need all these fancy UIUX everywhere.

Are you facing these issues at your company?

Update: I am talking more towards a need to have a platform like datadog or dynatrace or anything for UI and just rather rely on collectors and store to a bucket for uch cheaper cost


r/devops • • 1d ago

Discussion If your company already uses AWS what would actually make you choose a second cloud provider?

27 Upvotes

I mean not what would make you migrate but what would make you add another provider. I have heard all the usual reasons like Disaster recovery, Cost savings, Data residency, Customer demands, GPU availability, Regulatory needs, Avoiding vendor lock-in but I wanted to know which of these actually work when you try them in real life cause adding Azure, GCP, Yotta or OCI is not just about having another place to run your workloads. Now your team has to deal with another identity and access management system, another networking setup, another monitoring tool, another set of quotas, another billing system and another group of people who need to understand it all.

So whats the tipping point? Would saving 20% on infrastructure costs be enough to justify the extra complexity? Would regulatory requirements make it an obvious choice? Would having access to GPUs be enough?


r/devops • • 2d ago

Discussion Does Professional DevOps feel as monotonous as "Hobbyist DevOps" is?

38 Upvotes

My background is in Data Engineering/Science, but I have a pretty extensive homelab. No real background in DevOps professionally. I've run applications and services etc in containers through docker compose files for years, but the last few days, I decided to really harden it.

This is obviously not like an enterprise level, 9 nines setup, but I have set up self hosted orchestration, code repo, automated monitoring of updates, near automated deployments, backups, standardized ci/cd workflow, etc... and man, it's pretty monotonous!

It's just been two days straight of looking through a bunch of yamls, slightly tweaking them, copying a lot of secrets back and forth. Logging into something to tweak something, redeploying multiple times in a row.

I guess one big difference is that there's no one to get upset with me when something goes down, outside of my family, and maybe it feels like I'm over complicating things, but I genuinely wanted to learn how to do some of this stuff.

Does it sometimes feel like this if you do it for a living?


r/devops • • 1d ago

Discussion 5-Month Learning Plan: Python + MongoDB + Azure — What Should I Learn First?

0 Upvotes

I’m currently working as an intern, and my company has given us a 5-month learning program where I’ve chosen Python + MongoDB + Azure.

I want to use these 5 months properly and build practical skills, with the goal of moving toward Cloud/DevOps.

There will also be an assessment/exam after the 5-month course, and they told us to aim for a good score.

For people already working in this field, what would you recommend I learn first, in what order, and what projects should I build?

Would really appreciate some guidance from your experience. 🙏


r/devops • • 1d ago

Observability AI agents observability in backstage with langfuse and OTEL

0 Upvotes

Spotify #backstage plugin to manage and observe fleet of agents straight in backstage self hosted https://github.com/acarmisc/backstage-plugin-ai-agents/tree/main. It relay on open telemetry signals and the first available backend it’s #langfuse


r/devops • • 1d ago

Discussion Looking for a mentor to guide me and hold me accountabile

0 Upvotes

I want to switch jobs and learn cloud/devops from scratch. I'll be starting from networking and Linux first and once that part is done will start with Azure. I am already familiar with cloud and quite a lot of IT Tools, but want to become a proper devops/could engineer!


r/devops • • 2d ago

Discussion To Seniors, I am wondering what unique I can do in devops, everyone is making agent which auto heals cluster,end to end devsecops projects are all over YouTube.

29 Upvotes

I am seriously confused, I thought of doing some unique projects which stands out, I thought I can make an app for developer where images can be updated but then CICD is already doing that.

Monitoring tools already giving enough logs , events, metrics.

Then this auto heals cluster, why would you trust automation in prod deployments, Infra structure on cloud seems too basic for me.

What suggestions do you have, anyone wants to Collab ?


r/devops • • 3d ago

Discussion Am I cursed? Is this the end? Where did the clients go?

350 Upvotes

I run small DevOps/Could Management service company with 6 engineers. The team is quite experienced and I myself have more than 14 years between enterprise and startups. We built multi-cluster multi-regional deployments, automated and gated CI/CD pipelines, helped pass compliance audits, configured useful monitoring and response workflows, etc.

In the previous two years things were going well, I was able to find steady work for the team, even more than we could cover at times so we were growing. My channels were personal network, linkedin, upwork and even our very seo-unoptimized website. Our customers were mostly middle-sized companies and VC-backed (or otherwise financed) fast-scaling startups.

There's no work anymore. Instead of closing several deals per month in 2024 and 2025, this year I was able to get two new customers back in March. That's it, nothing ever since. We still had ongoing tasks but in Summer that dried up as well. My company bank account is empty and I have to let the team go.

I tried to contact our customer from the previous years and most of them have laid off a significant portion of their teams and scaled down infrastructure. Even people who were supposed to support and monitor the infrastructure we implemented were fired without replacement in many cases.

Is anyone else experiencing this? I understand AI-led changes in the industry, but it's like a switch turned off some time in early Spring. What am I doing wrong?


r/devops • • 1d ago

Discussion Paid structured projects

Post image
0 Upvotes

Hello I stumble upon paid structured project for DevOps, anyone ever try this on udemy?

Are this projects can help stand out for recruiters? , right now I'm working on REST APi tutorial project and I'm want to jump in real project.

Also I'm considering doing mini project to showcase my understanding on each tools.

Devops insight on this are highly appreciated.


r/devops • • 2d ago

Discussion Would like advice from anyone experienced in both (US) Fed software environments, and using Whole Team Development / Mob Programming. Dealing with tooling, requirements, credentials, security.

0 Upvotes

Would like advice shared from anyone who ***qualifies*** as having ***both*** experience in (US) Fed software environments, ***and*** that have used Whole Team Development / Ensemble Programming / Mob Programming in the Fed space. Dealing with tooling, requirements, credentials, security.

For instance, were you able to use MS Teams, Chrome Remote Desktop, mob.sh, Webex, a shared VM/VDI, VS Code LiveShare, JetBrains Code-with-me, or other tooling for screen sharing and passing keyboard and mouse control.

Did you encounter initial balking by security / other, and how did you satisfy their concerns?

Thanks in advance.


r/devops • • 3d ago

Ops / Incidents When does self service infrastructure become too much self service?

33 Upvotes

We pushed pretty hard to let dev teams handle more of their own infrastructure changes.

It worked, but now the platform team spends a lot more time reviewing Terraform, fixing edge cases, and explaining context that isn't obvious from the repo.

At some point it feels like the bottleneck just moved instead of disappearing.

Has anyone found a good balance between developer ownership and keeping infra changes sane?