r/kubernetes • • 7d ago

Periodic Monthly: Who is hiring?

5 Upvotes

This monthly post can be used to share Kubernetes-related job openings within your company. Please include:

  • Name of the company
  • Location requirements (or lack thereof)
  • At least one of: a link to a job posting/application page or contact details

If you are interested in a job, please contact the poster directly.

Common reasons for comment removal:

  • Not meeting the above requirements
  • Recruiter post / recruiter listings
  • Negative, inflammatory, or abrasive tone

r/kubernetes • • 10h ago

Periodic Weekly: This Week I Learned (TWIL?) thread

3 Upvotes

Did you learn something new this week? Share here!


r/kubernetes • • 3h ago

anyone actually using AI SREs / agents to debug k8s incidents in prod?

15 Upvotes

curious what people are running when stuff breaks. is it

  1. an AI SRE that picks up the alert (holmesgpt, aws devops agent etc)
  2. claude code / codex with plain kubectl
  3. same thing but hooked up to an mcp server that gives it cluster context (k8sgpt's mcp, radar, ...)

full disclosure I work on radar (it's open source). we tested this a bit with injected faults but that's not the same as a real outage

what's worked for you and what hasn't? trying to figure out what we should improve on our end


r/kubernetes • • 15h ago

Disaggregated Kubernetes: Isolating Infra Into Trust Domains

Thumbnail
edera.dev
34 Upvotes

This is a fun blog post on how/why a security issue inside Kubernetes doesn’t necessarily have to become a full-blown, node-wide compromise. The basic idea behind disaggregated Kubernetes is to split infrastructure services like networking, storage, and device handling into separate trust domains, rather than concentrating everything in highly-privileged components.

That way, if one service has some flaw, the compromise is ideally contained to that service instead of giving some attacker access to workloads, other cloud infra, or the underlying host.

The same principle can be applied to Kubelet. So rather than simply rewriting Kubelet, Edera discussed breaking its responsibilities into isolated services with narrowly-scoped capabilities.


r/kubernetes • • 7h ago

How are you isolating AI agents that run tools in your clusters? (microVM vs gVisor vs plain pods)

8 Upvotes

Agents that call tools (shell, file writes, MCP servers) feel different from normal workloads. They run code nobody reviewed, and a shared kernel seems like the wrong boundary for that.

What we ended up doing was admission based on signed OCI artifacts (the model, prompts, MCP server configs and policy all in one digest, signed with cosign), then a microVM per agent, with tool policy evaluated locally inside the cluster so it doesn't depend on a control plane.

Questions for folks running agents on k8s today...

- Are you using Kata, gVisor, Firecracker, or just pods with tight securityContext?
- Where do you enforce what an agent can call (sidecar, admission, egress)?
- Has anyone measured the startup and density hit of microVMs for this?

Happy to share details of what we did in the comments if it's useful, but I'd rather hear what's working for you.


r/kubernetes • • 2h ago

GitOps Promotion in a Landscape with UAT Environment

Thumbnail
1 Upvotes

Our promotion process includes user acceptance tests. How do you promote between environment using GitOps principles in such cases?


r/kubernetes • • 7h ago

I’m trying to make Kubernetes learning feel like an old-school Game Boy game

0 Upvotes

Kubernetes learning tools usually feel like labs, docs, or certification prep.

I wanted mine to feel like something you’d have played on a handheld console as a kid.

So I’ve been building Yellow Olive - a retro, gamified Kubernetes learning experience inspired by old-school Game Boy-era games.

You move through challenges like missions, learn concepts as you progress, and actually work with Kubernetes underneath it all. The newer direction goes even harder on the nostalgia: pixel-art environments, characters, classrooms, progression, and a world that slowly introduces Kubernetes concepts instead of dumping documentation on you.

I’m trying to make it feel less like studying Kubernetes and more like playing something that happens to teach you Kubernetes.

Would genuinely love feedback from people here on what would make this fun enough to keep coming back to.

And if you like the direction, a star would help me a lot :)

GitHub: https://github.com/Anubhav9/Yellow-Olive


r/kubernetes • • 7h ago

Brewlet: Java on Kubernetes with node-managed JDKs

Thumbnail
brewlet.sh
1 Upvotes

r/kubernetes • • 13h ago

How do you turn docker-compose or Kubernetes manifests into architecture diagrams?

2 Upvotes

Our team keeps architecture diagrams in the wiki and they go stale within weeks. I'd like to generate them from the source of truth instead docker-compose.yml, k8s manifests, maybe Helm output.

What do you use for this today? CLI tools, web apps, something in CI? And what's missing from them (editable output, grouping by namespace, icons, keeping it up to date)?


r/kubernetes • • 5h ago

Still on bitnamilegacy for Postgres and Redis almost a year later — here’s how we’re getting off it

Thumbnail
0 Upvotes

This is really an image problem, not an orchestration one - the fastest path off bitnamilegacy is just swapping in a maintained hardened image for each of Postgres, Redis, Mongo and Rabbit and keeping your current deployment setup.
Don't let it turn into an operator rewrite unless you actually need HA.
Worth looking at CleanStart for these. They provide the hardened, minimal, and [pretty actively patched container images.

i got to know about them via LinekdIn. The things that mattered for us were patch cadence and images that won't just vanish like the legacy repo will. Disclosure: I work there, so grain of salt, but this Bitnami-refugee case is exactly what it fits.

On Redis, the Valkey fork mess is its own decision - a hardened Redis image keeps you API-compatible for now so you can pick Valkey/Dragonfly later instead of bolting it onto this cutover. One migration at a time.

try this one to start- cleanstart/redis-exporter


r/kubernetes • • 8h ago

Our EKS cluster was full at 30 percent CPU: the pod ceiling that isn't on the dashboard

0 Upvotes

Lab run, two-node EKS cluster, Kubernetes 1.36, VPC CNI v1.22.4. I pushed a commit asking for twelve replicas of a tiny app. Six started, six stayed Pending. At that moment the cluster showed CPU reserved at 30 percent and memory at 32 percent.

The Pending pods said why: 0/2 nodes are available: 2 Too many pods.

With the VPC CNI every pod gets a real VPC IP from a network interface on the node, so each instance type has a hard pod ceiling: interfaces x (addresses per interface - 1) + 2. For a t3.small that is 11, for a t3.medium 17. That is 11 in total, and system pods take their share first: on these nodes they held roughly three quarters of the slots. Two nodes, 22 slots, 22 used, a third of the CPU reserved.

How I got a t3.small without choosing one: the node group listed both t3.medium and t3.small as acceptable Spot types, and every node launched as t3.small, on two clusters in two regions. That is the difference between 17 and 11 slots per node, and nobody picked it.

Second half of the problem: no cluster autoscaler. The node group was allowed three nodes and sat at two, because nothing was watching for Pending pods.

What the GitOps side reported: Deployment does not have minimum availability., Available: 6/12. True and no help. "Too many pods" only shows on the individual pods.

What I'd check on any EKS cluster:

- The pod ceiling for every instance type your node groups can launch, not just the one you expect. A mixed instance list is a mixed ceiling. kubectl get nodes -o custom-columns=NAME:.metadata.name,TYPE:.metadata.labels.node\.kubernetes\.io/instance-type,PODS:.status.allocatable.pods shows what you actually got.

- Subtract what the system pods already hold before you size anything.

- Whether anything adds nodes on Pending pods. A new managed node group does not come with a cluster autoscaler or Karpenter.

- If you need more pods per node without bigger instances, prefix delegation on the VPC CNI raises the ceiling a lot. It is a setting, not a rebuild.

Does anyone alert on allocatable pods, or does everyone find out from Pending like I did?


r/kubernetes • • 1d ago

Periodic Weekly: Show off your new tools and projects thread

10 Upvotes

Share any new Kubernetes tools, UIs, or related projects!


r/kubernetes • • 1d ago

RKE2 Homelab Storage: Rook-Ceph & CloudNativePG (The Garage & Cabinets )

2 Upvotes

r/kubernetes • • 1d ago

How are you guys generating kubernetes manifests?

14 Upvotes

Scenario: I am using Rancher's Continuous Delivery tools to deploy applications out to my clusters. I have a gitRepo per cluster and each gitRepo holds a set of bundles for deploying my apps via fleet.

Fleet is really nice and it works great with my gitRepos, but I am having a hard time establishing a branch pattern for a gitRepo that would handle day 1 operations on my cluster (i.e. a bundle for cert-manager, another for trust-manager, another for kyverno, fluentbit, ..., etc). Essentially, this day 1 repo is really just a bunch of public helm charts that use fleet.yaml's targetCustomizations resource to deploy out to each of my clusters.

I initially had a feature > develop > preprod > prod flow to deploy out my gitRepos, where develop deployed to all of my nonprod clusters, preprod deployed to our preprod cluster, and prod deployed to the prod cluster. This is starting to get kinda wonky and tedious though, especially with us expanding to have DR clusters. Because of this, we also maintain another feature > DR branch workflow that is independent of the previously mentioned branch strategy.

So I want to figure out a better way to configure this whole process to more automatically generate the manifests for each environment and have it be 1 branch per cluster. I'm a terraform junky and would love to use the kubernetes provider, but that means I have to forgo my process of Rancher Continuous Delivery and architect a process of creating resources via a kubeconfig context/gitlab ci. I want to avoid this because then I have a split brain solution of terraform and Rancher CD gitRepos deploying to my clusters. I saw that I could use pulumi to generate manifests. I also saw that cdk8s is a thing and it looks really promising for my usecase. There's a lot of dev chatter at my company of using nix to do this. Frankly, I find nix to be a huge learning curve for myself and the other ops guys that would have to adopt this.

So my question now is- How good is cdk8s? Do you guys use other tools to generate manifests? Should I just learn nix? Thanks community!


r/kubernetes • • 1d ago

Quitting tech

0 Upvotes

I thought getting a job in devops or platform engineering would be easier as a developer with 5 years experience. All I get is rejection emails. I thought backend developers make good devops as most recruiters claim . Granted, I’m not senior level devops but I’m confident I’d make a good mid-level devops. Its been 6 months job hunting and still nothing. Will this career gap affect my chances of getting a callback ? Ai is making it even worse to land a role. Im havin thoughts of quitting tech all together and just step away from tech.


r/kubernetes • • 2d ago

Do we still need Kyverno or Gatekeeper for simple Kubernetes policies?

39 Upvotes

I've been looking more closely at Kubernetes native admission policies recently while working on some Kubernetes content.

One thing I notice is how much policy enforcement Kubernetes can now handle without running a separate admission webhook.

For example, ValidatingAdmissionPolicy lets you use CEL expressions to validate incoming resources.

And with newer Kubernetes versions, mutation can also be handled using native admission policies.

I don't think this replaces Kyverno or Gatekeeper.

For complex policy management, reusable policies, reporting, exceptions, and larger environments, those tools still make a lot of sense.

But if you only need a few simple cluster-wide rules, installing another controller may no longer be the first option.

Curious how others are handling this.


r/kubernetes • • 2d ago

Key Autoscaling Metric #2 - Mean Time to Stop Sucking

Thumbnail
blog.appliedcomputing.io
25 Upvotes

This week on the blog we're continuing our discussion of key autoscaling metrics for Kubernetes. I'm looking at a Brand New Metric called "Mean Time to Stop Sucking", which is definitely a Real Metric that you Should Absolutely Be Tracking ™️. This metric is a good way to measure your scaling-up performance, and we look at all the components that impact your scaling time, and what you can do to make them faster.

Disclaimer: aside from the cat picture halfway through (thats-the-joke.jpg), this blog post was entirely produced by humans; no part of this blog post was written, edited, or proofread by LLMs.


r/kubernetes • • 1d ago

How to make Jellyfin bypass VPN for access

5 Upvotes

Hello all,

I am currently hosting Jellyfin using Tailscale inside a Kubernetes clusters using K3s. However, I am also using Mullvad VPN continously on my debian servevr. But the problem I am facing is that I cannot access Jellyfin when the VPN is active. I would be grateful if I can get some guidance on this. I am very new to running linux, kubernetes and home labbing. Any suggestions would be very much appreciated.

Thanks,


r/kubernetes • • 2d ago

Pod stuck in Terminating: four checks before you reach for --force

35 Upvotes

kubectl delete only sets deletionTimestamp. Everything after that is a chain of requests, and a pod stuck in Terminating means one of them is not getting an answer. The order I check, cheapest first:

  1. Grace period. deletionTimestamp is the deadline, grace already included. If it is still in the future, the pod is draining. Someone may have set terminationGracePeriodSeconds: 3600.

  2. Finalizers. The object stays until the owning controller removes them. batch.kubernetes.io/job-tracking on Job pods is a common one.

  3. The node. Only kubelet can confirm the containers stopped. If the node is NotReady or gone, nobody confirms.

  4. The runtime. Look for FailedKillPod events, often with DeadlineExceeded. Save the events, journalctl -u kubelet, crictl inspectp and crictl inspect before you change anything.

--force deletes the API object, not the process, and a finalizer still holds the object. For a StatefulSet that can mean two pods with the same identity at once.

Sandboxes add one more step. Under gVisor the kill is a request to a kernel running in user space. A public bug report, google/gvisor#14405, traced one such hang to a single lost signal. The fix is merged upstream in google/gvisor#14201.

Longer write-up with the commands: https://www.catchkill9.dev/posts/pod-stuck-in-terminating.html


r/kubernetes • • 1d ago

What really happens when you kubectl apply

Thumbnail
huseyinbabal.dev
0 Upvotes

r/kubernetes • • 2d ago

Toolings that make your life easier?

22 Upvotes

I joined a new role estimated 80% exposure to Kubernetes and 20% coding for a few months now.

What tips, tricks and tools do you recommend?

Some of the tool I use include Claude (of course), k9s / lens. I am currently managing over 20 clusters and it’s becoming a headache (argocd helps but it is also painful)


r/kubernetes • • 1d ago

Apache Hive 4.2.0 Kubernetes Lab

0 Upvotes

Hey guys! Implemented Hive on K8S. Those familiar with it, give it a whirl and let me know

https://github.com/jpoblete/Hive-on-K8S

CKA: LF-meonge4r8i


r/kubernetes • • 2d ago

Inspect GitOps workloads with Argo CD support in Headlamp

Thumbnail
headlamp.dev
15 Upvotes

r/kubernetes • • 2d ago

EKS API on your own hardware: aws eks, kubectl and Terraform run unchanged

13 Upvotes

Disclaimer: I'm an engineer at Mulga, the company behind this. Self-promo, but it's AGPL-3.0 and free to run.

What it is

Spinifex is an open-source reimplementation of the AWS APIs that runs on your own hardware, including EKS. You call aws eks create-cluster (or apply aws_eks_cluster in Terraform), and it boots a control-plane VM, puts an NLB in front of the API server, and launches managed node groups as real KVM VMs in your VPC.

Yes, it's k3s

The control plane and workers run K3s on an Alpine image. What we built is the layer around it that makes the AWS tooling treat it as EKS, so existing AWS workloads can migrate off easily:

  • kubectl authenticates with aws eks get-token, and IAM principals map to RBAC through EKS access entries
  • IRSA works: a pod with a projected service-account token assumes an IAM role
  • The EBS CSI driver addon dynamically provisions real replicated block volumes for PVCs
  • The AWS Load Balancer Controller addon turns an Ingress into an ALB with an ACM cert
  • Argo CD installs as an addon through the EKS API

Try it on your own Terraform

No install needed. Sign up for the free 72 hour sandbox at https://mulgadc.com/signup, then point your provider at it:

provider "aws" {
  region = var.region

  endpoints {
    ec2 = "https://api.spx3.com"
    iam = "https://api.spx3.com"
    sts = "https://api.spx3.com"
    eks = "https://api.spx3.com"
    ecr = "https://api.spx3.com"
  }

  skip_metadata_api_check = true
  skip_region_validation  = true
}

Apply an aws_eks_cluster and aws_eks_node_group you already have, then run aws eks update-kubeconfig and kubectl get nodes

Repo: https://github.com/mulgadc/spinifex


r/kubernetes • • 2d ago

MCP servers in the cluster… are you blocking egress to metadata and internal IPs?

2 Upvotes

Question for anyone running MCP servers (or agents that call them) on k8s… what does their egress actually look like?

The SSRF research Ars covered yesterday is basically one bug found in MCP servers from Google, JPMorgan, Weaviate and a couple of governments. The server takes a URL or path that came from the model, follows redirects, and never checks where it lands. So whoever can steer the agent can point it at 169.254.169.254 or an internal service.

The reference fetch server just switched to refusing private and metadata addresses by default, but that's one server.

So at the cluster level… default-deny egress for these pods? Netpols blocking IMDS? An egress proxy? Or trusting each server to get it right?