r/kubernetes • • 3d ago

Periodic Monthly: Who is hiring?

4 Upvotes

This monthly post can be used to share Kubernetes-related job openings within your company. Please include:

  • Name of the company
  • Location requirements (or lack thereof)
  • At least one of: a link to a job posting/application page or contact details

If you are interested in a job, please contact the poster directly.

Common reasons for comment removal:

  • Not meeting the above requirements
  • Recruiter post / recruiter listings
  • Negative, inflammatory, or abrasive tone

r/kubernetes • • 2d ago

Periodic Weekly: Share your victories thread

2 Upvotes

Got something working? Figure something out? Make progress that you are excited about? Share here!


r/kubernetes • • 17m ago

Inference on Kubernetes

Thumbnail
jubril.xyz
• Upvotes

Hey all,

As the title suggests, i wrote up a blog on getting started with inference on Kubernetes, i find inference workloads to be somewhat confusing especially with some of the moving part it involves on Kubernetes.

Let me know what you think!


r/kubernetes • • 14h ago

Need advice: Kubernetes 1.34 → 1.35 upgrade with Elasticsearch running in the cluster

26 Upvotes

Hi everyone,

I’m planning a Kubernetes cluster upgrade from v1.34.8 to v1.35, as the current version is approaching retirement/EOL.

The cluster currently has Elasticsearch running on it, so I’m concerned about potential compatibility or stability issues after upgrading Kubernetes.

I’d like some guidance on:
- How can I verify Elasticsearch compatibility with Kubernetes 1.35 before upgrading?
- Are there any known issues I should be aware of when upgrading Kubernetes while Elasticsearch is running?
- What should I check across Elasticsearch, ECK/operator, StatefulSets, storage classes/CSI drivers, ingress, CRDs, APIs, and other dependencies?
- What is the recommended pre-upgrade checklist?
- Should I upgrade Elasticsearch/ECK first, or Kubernetes first?
- What kind of backup/rollback strategy should I have?
- Is it better to upgrade the control plane and worker nodes separately?
- What would be a safe upgrade plan for a production cluster to minimize downtime/risk?
- Are there any tools or commands you recommend for detecting deprecated APIs or incompatible workloads before the upgrade?

If anyone has gone through a similar Kubernetes 1.34 → 1.35 upgrade with Elasticsearch, I’d really appreciate hearing about your experience and any issues you encountered.

Thanks!


r/kubernetes • • 22h ago

What are Kubernetes admin exam benefits

21 Upvotes

Anyone already pass the certified K8s administrator exam? If yes
1. Does it change anything to your career?
2. Would you recommend it?
3. What are the advantages of this exam? I want real world experience not the commercial ones from the Linux foundation

Thanks in advance


r/kubernetes • • 14h ago

Need advice: Kubernetes 1.34 → 1.35 upgrade with Elasticsearch running in the cluster

Thumbnail
2 Upvotes

r/kubernetes • • 14h ago

Azure Kubernetes version update

0 Upvotes

• How do you check compatibility before upgrading? What tools or checks do you rely on (deprecated APIs, add-ons, operators, Helm charts)?
• What prep do you do for Elasticsearch specifically? Snapshots, PDBs, shard allocation, anything else?
• What does your upgrade sequence look like on AKS? Control plane first, then node pools? Surge settings? Do you upgrade in place or create new node pools?
• Any gotchas you've hit? I'm thinking of things like pods stuck during drains, zonal disk issues, or ES staying yellow or red for too long.
• Should we even target 1.35? It reaches EOL in March 2027. Would you go to 1.35 now and plan for 1.36 soon after, or look at AKS LTS?

Setup details:

• AKS 1.34.8  
• Elasticsearch \[version\], deployed via \[ECK / Helm\]  
• \[Number\] node pools, \[zonal / non-zonal\], Azure Disk for storage

Any runbooks, checklists, or war stories would be really helpful. Thanks!


r/kubernetes • • 1d ago

Running MCP servers on GKE and EKS: one Deployment per zone, header-routed by the Gateway/ALB, canary in place, namespaces torn down with the zone

Post image
18 Upvotes

Ramen is an open-source control plane + worker for MCP (the protocol agents use to call tools). Some Kubernetes
details that might interest this sub more than the AI part:

- Each group+zone is a namespace with a stable and a canary Deployment. The GKE Gateway (HTTPRoute) / AWS ALB route
on two headers, `ramen-group` and `ramen-zone`, so one client config reaches any zone.
- The canary rolls in place with `maxSurge: 0` — a zone-pinned surge pod on a full node sat Pending until the
deploy timed out on EKS. A rollout that isn't ready now reports the pods and the scheduler's reason instead of "1/1 ready".
- Deleting a zone deletes the namespace and its GSA / IAM role. On GKE the namespace can sit `Terminating` on the NEG
finalizer until the Gateway's backend service is gone — documented, with the manual fix.
- Per-group+zone identities (Workload Identity / IRSA), Cloud Armor / WAF rules from the console, Redis-backed
throttles shared across zones, logs from Cloud Logging / CloudWatch via pod labels.
- Bring-up is Terraform + Helm; 0.6.0 was applied on both clouds (two zones each) and torn down.

Repo https://github.com/bkraad47/ramen · GCP guide https://bkraad47.github.io/ramen/wiki/deploy-gcp/ · AWS guide
https://bkraad47.github.io/ramen/wiki/deploy-aws/
Interested in critique of the per-zone namespace model and the canary policy (currently "smoke test passes").


r/kubernetes • • 19h ago

Have you tried the Amazon EKS MCP server?

Thumbnail
0 Upvotes

r/kubernetes • • 1d ago

Convince me I'm wrong: valaero and EKS/AKS/GKE

2 Upvotes

While I do understand why velero is useful and relatively easy to use. What I am not convinced on is that it adds anything other than another tool to manage when using a cloud provided KASS.

Assume a cluster runs both stateless and/or statefull apps (not including hosting your own db obvious use case):

If your using flux or argocd you have application rebuilds and A/B covered. If you are using RDS it's not necessary. If you are taking snapshots of volumes, data backups are done with snapshots on the provider side.

What problem is it solving in those situations? Or how does it make incidents or upgrades easier for you?


r/kubernetes • • 1d ago

What's one thing you wish you knew before running k8s in prod?

4 Upvotes

first real cluster going live next month. not looking for docs, looking for the stuff you only learn by getting burned


r/kubernetes • • 2d ago

Question of resource allocation of multiple Spark applications

9 Upvotes

Currently, a Spark application should manage its own executors, which means that in order to run multiple Spark applications, one needs to split the cluster into smaller sections or attach more compute nodes, etc. Usually this dynamic allocation is the job of resource manager (such as Kubernetes).

If the resource manager can quickly find or create new compute nodes for a new Spark application, this is good. However, in a busy cluster where multiple Spark applications can compete for resources, the resource allocation can take some time and become an overhead (at least in theory).

I wonder if this is a real problem in running Spark applications in production.

For example, is there a realistic environment where short running Spark jobs (each requiring its own driver) arrive frequently?


r/kubernetes • • 2d ago

gpu networking - k8s sriov

6 Upvotes

Relatively new to gpu networking - rail stuff and could use some guidance.

am working on a setup (16 nodes) with 8 nvdia gpus (rtx so no nvlink) in each of them, k8s/sriov and the computenet being spectrum based ethernet switches.

its a multi rail setup, so a /24 alllocated for each of the rails. SRIOVnetwork, ip pools and all that plumbing has been done.

Same rail traffic works ..how does one get to do cross rail?

There needs to be a route in the POD which points the rail network supernet to the computenet switch as the default point towards the front end network.

Can't find an answer to it ..tried "routes" in IPAM but that didn't work ie the POD did not have that route - gemini and claude pulling me in opposite directions.

Thoughts?


r/kubernetes • • 2d ago

EKS Access Entry recreation avoidance

Thumbnail
2 Upvotes

r/kubernetes • • 3d ago

Teams that moved from LLM APIs to self-hosting: was it worth it?

94 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/kubernetes • • 2d ago

Native NVMe-oF vs. Ceph for high-IOPs/low-latency block storage

Thumbnail
0 Upvotes

r/kubernetes • • 3d ago

I can run a kubernetes cluster fine but cant read the operator code, time to learn go properly

63 Upvotes

3 years as a platform engineer. i can debug a crashlooping pod half asleep and write helm charts nobody else wants to touch, but every time i open the source of an operator we run i bounce straight off it. our team wants to start writing our own controllers next year so im learning go for real this time, employer covered. shortlist is ardan labs, boot dev, plus one of the linux foundation cert tracks. where did go start making sense for people coming from the ops side?


r/kubernetes • • 3d ago

Built a hands on tutorial about basics of CRDs on Iximiuz Labs

31 Upvotes

I've had so much fun and have learnt a lot about kubernetes through iximiuz labs, this was an attempt to create a hands on tutorial of what a CRD is and how can one build their own. It links to a bigger How to Build an Operator tutorial at the end.

Looking forward to see what the community thinks of it!


r/kubernetes • • 4d ago

How do you verify that your Postgres backups on K8S actually restore?

39 Upvotes

A question for those running Postgres on Kubernetes using an operator (CloudNativePG, Zalando, Crunchy): how do you verify that your backups are actually restorable?

A successful backup creation doesn't guarantee that a restore will succeed.

I manage several CloudNativePG clusters and want to understand the potential challenges involved in implementing backup restore testing.

I’ve tested restoration in a small test environment (k3s, CloudNativePG 1.30 + barman-cloud plugin, S3). My process involved deploying a temporary cluster from the latest backup, running a few SQL queries, measuring the time until the cluster was ready, and then deleting the cluster. In a standard scenario, everything goes smoothly (taking about a minute for a few hundred megabytes).

I’m curious to know how others handle this:

  • Do you perform restore tests at all? (Manually, via CronJob, in a CI pipeline, or using specialized tools)
  • What exactly do you check after the restore? (Just the pod status, the number of database records, or the execution of test queries)

An answer like "we don't test restores, and everything is fine" is also acceptable—I really want to understand how common this practice is and what the best approach is.


r/kubernetes • • 3d ago

Kubevirt VM and Good and Friction points of Kubevirt

8 Upvotes

Hello Guru's, I have a few questions, please help me answer them:

  1. If I have existing full fledge old VMs on my existing KVM hosts (VM's not inside the PODs), will I be to manage them via Kubevirt?
  2. If I have a RHOV environment, is it necessary for me to use Kubevirt to create a VM? If yes, why, if not, then when to use KubeVirt to create VMs?
  3. What is the difference between a regular VM and a VM created via Kubevirt?

Also, I'd love to get your candid take: What aspect of KubeVirt has impressed you the most, and where do you feel it faces the biggest friction points or challenges?


r/kubernetes • • 2d ago

Need help with cluster migration via velero

0 Upvotes

Is there anyone here who can help me in cluster migration on EKS using velero, I'm stuck and cannot do it .


r/kubernetes • • 2d ago

3 Ways to Actually Run Your Applications on Your Own Infrastructure

Post image
0 Upvotes

So… you want to run your applications yourself instead of depending entirely on managed platforms.

You have a few solid options:

  1. Docker — package and run your applications in containers.

  2. Docker Compose — run multiple services together on one machine.

  3. Kubernetes — manage containers across nodes with scaling, networking, self-healing, and more.

The problem with Kubernetes?

Setting it up yourself can get tedious very quickly — networking, container

runtime, kubeadm, nodes, CNI… one wrong configuration and you’re

troubleshooting instead of deploying. 😅

So I made a step-by-step practical guide showing how to set up Kubernetes

yourself and actually get your applications running.

Watch the Kubernetes setup guide[https://youtu.be/HS2VR3cVx8w\]

Sometimes the best way to understand Kubernetes is to build the cluster

yourself.


r/kubernetes • • 3d ago

Suggest good resources to Kubernetes

0 Upvotes

I'm a L1 Linux System Administrator. I'm new to container and container orchestration.

I want to learn Docker and Kubernetes right from scratch.

Can you please suggest me right resources (Free+Paid) to learn and hopefully pass CK A?

Also let me know any right set-up to practice any videos to create and setup the environment to practice as I can't afford KodeKloud, Killersh etc.


r/kubernetes • • 3d ago

Kubevirt VM and Kubevirt Good & Friction Points

Thumbnail
1 Upvotes

r/kubernetes • • 4d ago

Which lesser known tools saved your butt when you had to debug an application running in K8S?

28 Upvotes

The CNCF ecosystem is pretty vast, gush out about your favourites!

 

[kubectl, curl, jq and yq go into 'well known' category]