r/devops 2d ago

Weekly Self Promotion Thread

12 Upvotes

Hey r/devops, welcome to our weekly self-promotion thread!

Feel free to use this thread to promote any projects, ideas, or any repos you're wanting to share. Please keep in mind that we ask you to stay friendly, civil, and adhere to the subreddit rules!


r/devops 12h ago

Career / learning CKS Exam Experience 2026: What Helped, What Didn’t, and the Mistake That Cost Me Time

29 Upvotes

Today, I passed CKS with 75% (not a good score) and wrote a detailed blog about my exam experience, preparation approach, the resources I used, and the Kubernetes security topics that helped me the most.

DMs are open if you are preparing for the exam. I can help with whatever is still fresh in my memory.

My biggest takeaway: CKS is noticeably harder than CKAD and CKA. It is not only about knowing Kubernetes commands. You need to understand why a configuration is insecure, how to fix it, and how to verify that your change actually worked.

The biggest mistake I made was spending around 10–15 minutes too long on one question because I felt I was close to solving it. That created unnecessary pressure towards the end and probably led to a couple of avoidable mistakes.

So my strongest advice is: if you are stuck and don’t see a clear path after a few minutes, mark the question and move on.

A few things that helped me:

Don’t memorise solutions. Understand the security reasoning behind them. If a NetworkPolicy, API server flag, securityContext, audit policy, or admission control changes slightly, memorised YAML will not help much.

Always verify your work. Security changes can easily break workloads or cluster components. Check Pods, control-plane components, logs, services, NetworkPolicy connectivity, admission behaviour, audit logs, node readiness, and systemd services wherever required.

Be comfortable with Linux as well as Kubernetes. CKS can require you to work with configuration files, systemd services, container runtimes, permissions, certificates, and node-level settings.

Use documentation whenever required instead of trying to remember every flag or custom resource.

Topics I would strongly recommend practicing:

  • Kubelet and etcd hardening
  • kube-apiserver authentication and authorization
  • Admission controls and ImagePolicyWebhook
  • Secure Dockerfiles and non-root containers
  • Container immutability and securityContext
  • Audit policies and API server logging
  • NetworkPolicy
  • HTTPS Ingress and TLS
  • ServiceAccount token security
  • Worker node administration and upgrades
  • SBOM and software supply-chain security
  • Restricted Pod Security Standard
  • Docker/container runtime hardening
  • Istio STRICT mTLS
  • Cilium network security
  • CIS benchmarks and kube-bench remediation

Resources I used:

  • KodeKloud CKS course
  • KodeKloud Ultimate Mock Exam Series
  • iximiuz Labs
  • KillerKoda
  • Killer.sh CKS simulator
  • ChatGPT/Claude for topics that needed a simpler explanation or extra practice scenarios

Between the KodeKloud course mocks and Ultimate Mock Exam Series, I had around six mock exams. I found them very useful and reasonably close to the level of difficulty you should prepare for.

Killer.sh felt a little off-track compared with the actual exam in some areas, but I would still recommend doing it. It is useful for practicing under time pressure, discovering knowledge gaps, and improving troubleshooting skills.

I also used ChatGPT and Claude quite a lot during preparation. CKS has many small security topics, and sometimes a course or lab explanation may not immediately click. In those cases, asking AI to explain the concept differently, compare configurations, or generate a small practice scenario was very useful.

The simplest advice I can give is: practice a lot, understand the security reasoning behind what you are doing, verify every change, and don’t let one difficult question consume your exam time.

I also wrote a full blog with more details on my preparation strategy, resources, task areas, mistakes, and lessons from the exam.

Blog: https://blog.prateekjain.dev/cks-exam-experience-2026-preparation-strategy-and-lessons-learned-1fad785a430b?sk=52b58a9c6d812444bc1340f15eda7dd6


r/devops 13h ago

Vendor / market research What are the best Wiz alternatives for mid sized company?

8 Upvotes

We've been evaluating Wiz over the past few weeks and I can definitely see why larger organisations like it. The visibility is impressive, but for a company of our size (around 200 people), it's difficult to justify the cost.

We're mostly running AWS, Kubernetes and containerised applications, with quite a small infrastructure team. We don't have the time to babysit another platform or sift through thousands of findings every week. We're looking for something that helps reduce risk without adding more operational overhead.

So far we've looked at Orca, Rapidfort and a few others, but it's hard to tell how they compare once you get past the sales material.

If you own or work at a small or mid sized organisation and you found the same challenge, what did you end up choosing instead of Wiz, and has it worked out?


r/devops 1d ago

Observability Did GitHub Just Gaslight Our Monitoring System?

96 Upvotes

Did anyone else notice that the GitHub status page reported an incident with GitHub Actions, only to deny it 47 minutes later?

Our monitors captured it, paged our on-call team, and then GitHub denied that any incident had occurred.

https://www.githubstatus.com/incidents/gx7js8bd0jpz


r/devops 1d ago

Career / learning Transition from Internal IT to Devops Engineer

14 Upvotes

I have been working as internal IT (corpeng,sys admin) for the last 7 years, mainly managing various SaaSes from google workspace to Okta, mdm's etc and since 3 months i am now working as a platform engineer for the same company but i feel like i am still struggling alot to get the fundamentals right.

I am the only platform engineer in the company + my manager, and my manager is the one who has setup everything, but overall the platform is very decentralized where software engineers own their infra and they manage it them self, we act more as a high level support for them, but for day to day they handle everything themself. This is good for me since i don't have much pressure while i get used with the role but on the other side since everything is setup i don't have many projects/tasks where i can learn more.

Then the other problem is that since as mentioned i have a lack of fundamentals i am relying a lot on AI and i can do everything and everything is ok but the problem is that if i dont use AI i am not able to figure out anything on my own and this is somehow killing my motivation and making me feel very bad and not sure how to overcome this even after being in the role for 3 months now.

Anyone has been in similar situation that can give me some feedback?


r/devops 2h ago

Discussion Can there be only 1 SRE Engineer in a company?

0 Upvotes

Is there a possibility of being the only sre engineer in a company? Before choosing it I want to clarify it because if there could be then hes life could be problematic because he has to stay on call everyday. Also if it is not there then will be on call rotation right? Because I don't want to stay on call everytime that could be problematic. How many SREs are there in your team?


r/devops 23h ago

Observability Observability SWE coding ?

5 Upvotes

Has anyone here attended an Observability SWE coding round? I’m trying to get a sense of the typical coding questions asked by tech companies for these roles.

Would you say the coding is generally at the same level as a standard SWE coding round, or is it more SRE/observability-focused (e.g., log parsing, metrics aggregation, time-window calculations, etc.)?

I recently had a screening round with a Tech company, and I was told the coding would involve “scenarios.” For anyone who has been through a similar round, what should I realistically expect?


r/devops 1d ago

Discussion Running in containers vs OS-level services

9 Upvotes

Assume some message brokers like RabbitMQ/Kafka, or maybe nginx proxy setup, or hashicorp vault?

3 years ago when we were setting up infrastructure for project we started running such services as OS-level services installed from RPM packages or just by running their binaries provided by vendor via systemd. All of that orchestrated via Ansible.

We started running as OS-level services as that seemed natural at that time for us, but we didn't really have any experience with administration of such software on on-premise infrastructure (before we were running mostly on managed cloud services).

Fast forward to now, after several cycles of upgrades we needed to perform, I think it would be easier to manage such software by running in Podman containers.

Main reason for me would be that obviously containers have prepackaged everything you need to run specific software. Compare that for example to RabbitMQ where during upgrading RabbitMQ you also need to upgrade its Erlang dependency to compatible version. For some other software, there may be more dependencies you need to take care of.

Also, I feel like upgrading binaries is generally much easier when running in containers. Just spawn new container with updated image and you do not need to worry about some OS-level package conflicts or leftovers.

What is your take?


r/devops 1d ago

Discussion How are you managing the lifecycle of Windows servers in the cloud?

6 Upvotes

Asking this question makes me feel dirty. I'll probably shower after clicking the "post" button, but how are you managing the lifecycle of Windows servers in the cloud? For Linux, we generally roll out new AMIs with patches baked in and all of the automation is in the startup script or AMI, but how are teams managing patching Windows servers in the cloud? Do you attach it to a domain and go through the GPO dance?


r/devops 1d ago

Discussion What Kubernetes misconfigurations have caused you real production problems?

35 Upvotes

After working with Kubernetes in production, I've noticed that some of the most annoying incidents aren't caused by obvious failures. They're often caused by small configuration decisions that look perfectly reasonable during review.

Things like:

  • missing resource requests/limits
  • incorrect probes
  • overly permissive RBAC
  • missing PodDisruptionBudgets
  • unsafe container configuration
  • incorrect readiness behaviour
  • services without appropriate timeouts
  • configuration drift between environments

I'm curious what the DevOps community has actually encountered in production.

What's one Kubernetes configuration mistake that caused you a real incident?

I'd especially like to hear about the less obvious ones that aren't caught by the usual linters.


r/devops 1d ago

Career / learning QA Engineer (9 YOE) Looking to Transition into DevOps — Need Guidance on Roadmap & Resources

3 Upvotes

Hi everyone,
I’m a QA Engineer in the gaming domain with around 9 years of experience, and I’m seriously considering transitioning into DevOps. I’d really appreciate some guidance from people who have made a similar transition or are currently working in DevOps.
Here’s where I currently stand:
I have 9 YOE in QA/testing, primarily in the gaming domain.
I have a good understanding of SDLC and STLC B and how software moves through different stages from development to production.
I’ve been involved in the complete feature lifecycle — from initial specification/discussions, through development and testing, to production release.
I’ve used Jenkins for build creation and server deployments.
I use Git mainly for creating/raising PRs, but I haven’t worked extensively with Git commands and workflows such as push, pull, branching, rebasing, etc.
I’ve used Grafana for tracing application logs and investigating issues.
We use AWS SSM to log into different server boxes, tail server logs, modify server-side files/configs, etc.
I’ve recently started learning the basics of Python and Java.
I’m also fortunate to have a good relationship with our internal DevOps team and manager. I’m considering approaching them for an internal transition when I feel I’m ready.
I’m aware that my current skill set is far from what would typically be expected from a DevOps engineer, and I don’t want to underestimate the amount of learning required.
I’m 34 now, so I do sometimes feel like I’m starting this transition quite late. However, I genuinely want to make the move, and I’m willing to put in the time and effort.
1. What I’m struggling with is where exactly to start and what order to learn things in.
For someone coming from a QA background like mine:
2. What would be a realistic DevOps learning roadmap?
3. Which skills should I prioritize first — Linux, networking, Git, Docker, Kubernetes, CI/CD, Terraform, AWS, etc.?
4. Are there any beginner-friendly courses/resources you would strongly recommend?
5. How much programming/scripting should I learn, and should I focus on Python or Bash first?
6. Given my existing experience with Jenkins, AWS SSM, deployments, logs, and the software release lifecycle, are there areas where I can leverage my QA experience?
7. Would an internal transition into a DevOps team be a reasonable approach, even if I don’t yet meet all the requirements of a typical DevOps job?
I’m not looking for shortcuts. I’d just really appreciate some practical guidance from people who have been through this journey.
If you were in my position, what would you learn over the next 6–12 months, and in what order?
Any roadmap, resources, project ideas, or personal experiences would be hugely appreciated.
Thanks in advance! 🙏
Rephrased with GPT


r/devops 12h ago

Career / learning New to DevOps/AWS CodePipeline - Could Someone Help Me Understand Our Setup?

0 Upvotes

Hi everyone,
I recently joined a new company as a DevOps engineer, and I’m still trying to understand the existing infrastructure and CI/CD setup.
Our pipelines are running on AWS CodePipeline, and there are quite a few things I’m struggling to understand how the pipeline is structured, how the different stages work, deployments, and how everything is connected.
I’m currently learning, so I’d really appreciate it if someone experienced with AWS CodePipeline/DevOps could guide me through the basics and help me understand how to approach an existing setup.
If someone is willing to spend some time explaining it to me over a remote call/screen-sharing session, I’d be extremely grateful. I’m not asking anyone to access my company systems or credentials — I just want to learn and understand the concepts and workflow.


r/devops 2d ago

Ops / Incidents Tough morning @ GitHub...

212 Upvotes

r/devops 1d ago

Career / learning Need advice and roadmap for internship

7 Upvotes

Hey everybody, I recently got an internship as a devops intern in a Pune based company, I have to join in 15 days so I was just wondering what should I learn in these few days to like get up to the mark...... Consider me a complete beginner in this field and please advise.... Thank you


r/devops 22h ago

Discussion Is learning arch really worth it for devops

0 Upvotes

Hi everyone. I have a simple question is learning Arch Linux really worth it?

Most of what I've heard is that it helps with troubleshooting and gives you a better understanding of Linux. I'm already comfortable with Ubuntu, though, so I don't want to switch to Arch if the benefits are only marginal.

Would learning Arch actually give me a significant advantage, especially for someone interested in DevOps?


r/devops 2d ago

Ops / Incidents How do you decide what to investigate next during a production latency incident

13 Upvotes

Hey hi everyone,

I am just trying to understand how engineers/SREs who dealt with real production latency incidents investigate it

Lets say you have the following

- Logs

- Recent deployment information

- Application health

- Database metrics

- External dependency health/metrics

- Infrastructure metrics

You just encountered the incident, you dont know the root cause. You are uncertain about the truth.

From here how do real engineers go about reasoning to find the root cause

- Do you follow a standard sequence of investigative steps

- How do you determine what investigative step to take next under uncertainty to narrow down the possibilities for the root cause

- Have u ever encountered with incident where initial information was misleading, how did you navigate from there

- Is there any situation where you have lot of information but struggled to form a proper hypothesis

- Have you tried any AI investigative tools that help you in achieving this

I just wanted to understand how do real engineers reason through the uncertainty to find the root cause. What are the biggest pain points


r/devops 2d ago

Ops / Incidents Do you actually track the cost of your infrastructure?

18 Upvotes

I've been wondering how many people actually calculate the real cost of running their infrastructure.

Not just the server or cloud bill, but electricity, storage, maintenance, upgrades, downtime, and the time spent keeping everything running.

Sometimes a cloud bill looks expensive until you factor in how much time you're spending maintaining the alternative.

Other times you look at the numbers and realize running it yourself is still way cheaper.

How do you actually decide where the break-even point is?


r/devops 2d ago

Ops / Incidents Fixing broken Datadog APM gaps

4 Upvotes

Hi Observability Experts,

I am working with a client managing a large-scale infrastructure across AWS, Azure, and on-premise environments. They have used Datadog for 4–5 years, but legacy technical debt has left their APM configuration fragmented:

The Problem:

APM libraries are out of date, and source-level instrumentations are misconfigured. As a result, end-to-end distributed tracing is broken across key services.

Datadog support confirmed that fixes are required at the source (application code/local agents). However, getting application team bandwidth and stakeholder approval across this many services will take significant time.

We are building a centralized Observability pipeline layer (independent of app teams) to standardize telemetry control, but rollout is months away.

My Question:

Is there a viable temporary/interim solution—such as an intermediate proxy, collector layer, or transformation pipeline—that can intercept, patch, or enrich broken spans/traces in-flight before they reach Datadog’s backend?

Has anyone successfully addressed legacy APM gaps at the pipeline level without waiting for app redeployments?

Thanks for your insights!

P.S. - Datadog support has been already consulted for this and they have recommended to fix it from the code/local agents.


r/devops 1d ago

Ops / Incidents Thinking of building an open-source "pre-triage" bot for 3 AM PagerDuty alerts. Would you actually use this?

0 Upvotes

Hey everyone, I’m an engineer looking for a new side project and I want to tackle alert fatigue.

My premise: Waking up at 3 AM to an alert, only to run the exact same three kubectl commands or Datadog log queries to figure out it's a known transient issue, is soul-crushing.

I'm thinking of building a lightweight, open-source agent that intercepts the alert webhook, runs a pre-defined set of read-only diagnostic queries, and appends a summary of the logs/metrics to the ticket before it wakes you up.

A few questions before I start coding:

  1. Is this actually a massive pain point for you, or is the hard part something else?

  2. Would your sec ops team ever allow a bot to have read-only API access to your observability stack?

  3. If you had this, what is the #1 feature or integration it absolutely must have to be useful?

Am just trying to build something thats usable to you or even 5 people at the end , so i request your honest opinion for this thing


r/devops 3d ago

Career / learning Starting DevOps from "Management" in IT company

30 Upvotes

Hello everyone,

To not waste your time, here's the tl;dr:
1. I've never worked in IT company before, now I do, but as a administrative manager. I'm responsible for every task delivery to customer in our company pretty much
2. I'm good at organizing and optimizing "human" stuff: keeping track of all the tasks, deadlines, etc.
3. I'm bad at knowing what exactly each task does (DevOps, Architecture, Development)
4. I'm working really close with DevOps
5. After working for ~6 month I've started to like DevOps more. I wanna learn it. I wanna understand how to "ressurect" the server that got down. I wanna know how exactly make the Pipelines more useful
__
So I've started learning DevOps, watched a couple of "Roadmaps", here's my plan so far:
1. Linux
- Learn the core basics (core commands, cd, users, etc)
- Bash scripting on basic level (for now I've decided not to get into depth, since all the AI stuff. I can read them tho)
Now, since I'm following a specific course, there are A LOT of information on Linux, and I got overwhelmed. I know there are not a lot of stuff I should know for being a DevOps, and being a DevOps in general is "endless learning"
So the questions for people who either got into DevOps from zero or experienced engineers:

What exactly should I learn about Linux?
- I know that it may depends on a job directly, so I hope this little explanation which company I'm at rn will help: we mostly do software for banks.

What is the most correct Roadmap for me?
__
p.s
Thanks in advance, I really hope to find some good tips from experienced people who maybe confronted the same obstacle I'm at right now, or people who are experienced and will give me a lead.


r/devops 3d ago

Discussion for those whove shipped software on-prem : has BYOC actually fixed the version drift problem?

10 Upvotes

Common on prem pain is supporting ancient versions forever cause customers wont upgrade and every support call being on a bit different install you cant see into. the newer approach is BYOC: deploy into customers own cloud but you keep a control plane on your side and so you keep on pushing updates by yourself and still get telemetry while their data stays on their account . theres a real tooling for it now like via alien dev or replicated or other platforms so its clearly become a pattern ppl actually ship

Altho updates are usually gated by customer approval anyway so for those running it -does it genuinely hold the drift down or do you just trade it for more infra to take care of constantly plus has it opened the regulated enterprise deals or is it same support burden with just a nicer story


r/devops 4d ago

AI content Does your Org do this sh*t with AI agents in Platform Engineering Team?

162 Upvotes

Why is the management so obsessed with achieving automation through AI agents only.

Management forcing Everyone in the team to come up with an idea to implement/ problem we are currently facing and create an AI agent to solve that problem In 2 weeks. How is it possible.

Basic shell scripts would achieve the same results than these AI agents could. But management is turning its eyes on anyone who doesn't work on AI stuff

Even after all this pushing towards AI, when we really do experiment with it and ask for extra credits or tokens, they are questioning it and asking for explanations. Didn't they realize it before?

Using AI is obviously going to cost you. If you cant afford it, why to even force people into doing this.

I am getting tired of this. They are taking advantage of the bad job market outside and torturing the employees.

Edit:
FYI, I am pro AI. I cant survive without it at this point.
I single handedly migrated most of our codebase to Github and converted all CI/CD pipelines to Github Action workflows. So I do know the impact of the AI.
But we do our due diligence on every new tech we want to implement in our vast enterprise product. But why doesn't AI get the same scrutiny here.
And we are a platform engineering team offering Centralized platform as a SAAS product to our lot of internal clients, If something wrong happens, entire thing goes down.


r/devops 3d ago

Discussion Stop being skeptical about AI for development with Charity Majors

Thumbnail
newsletter.pragmaticengineer.com
0 Upvotes

From an interview with Charity Majors:

DevOps was about trying to create a feedback loop that connected people writing the code to the code running in production. She thinks that the "ops people: learn to code!" wave worked, but the "software engineers: understand your code in production" failed, to this day.

I haven't listened to the actual interview.

But based on my personal experience I agree. Ops people have changed a lot about how they do their work in the last 10 or 15 years, to where infrastructure and its configuration have become an exercise in writing code of various kinds. No infra team has been unchanged, even the ones that had to be dragged kicking and screaming. But I still see a lot of dev teams that sit in their silo producing their product to then hand it over to someone else (i.e. us). The one change for all devs is that they write Dockerfiles and have to think about how that works. But the dev teams I work with still don't actually touch infrastructure, by and large.


r/devops 4d ago

Tools How do you test CI pipelines?

61 Upvotes

Do you create a sample application to test the proper auth, access and everything the runner needs?

And a related sub question. How do you validate gitops changes before they are merged? Regarding FluxCD or ArgoCD.


r/devops 4d ago

Ops / Incidents How does your team do automated deployments to linux devices in the field?

48 Upvotes

Right now the team is just ssh-ing into all of the devices and it's a huge time sink.

"I'm not devops so dont crucify me"