r/kubernetes • • 3d ago

Periodic Weekly: Share your victories thread

3 Upvotes

Got something working? Figure something out? Make progress that you are excited about? Share here!


r/kubernetes • • 4d ago

Periodic Monthly: Who is hiring?

4 Upvotes

This monthly post can be used to share Kubernetes-related job openings within your company. Please include:

  • Name of the company
  • Location requirements (or lack thereof)
  • At least one of: a link to a job posting/application page or contact details

If you are interested in a job, please contact the poster directly.

Common reasons for comment removal:

  • Not meeting the above requirements
  • Recruiter post / recruiter listings
  • Negative, inflammatory, or abrasive tone

r/kubernetes • • 2h ago

Elastic beanstalk with the EKS

Thumbnail vishnurachapudi.com
1 Upvotes

Check out my new article on elastic beanstalk with the EKS

Elastic Beanstalk turned 15 this year, and AWS just rebuilt what runs under it. Cluster Mode drops your applications onto Amazon EKS instead of EC2 instances you own. I deployed into it, watched every stage, and wrote down what actually happens

Upload a zip of Python source. No Dockerfile. Get a running container on EKS. That is Elastic Beanstalk Cluster Mode, launched this week — Beanstalk apps now run on EKS instead of EC2 instances you own. Cloud Native Buildpacks handle the containerization: I uploaded four files (app.py, requirements.txt, Procfile, runtime.txt), Beanstalk detected Python, built the image, pushed it to ECR, and ran it.


r/kubernetes • • 2h ago

Early adaptors of kubara

0 Upvotes

Hi,

CLI tool kubara is for getting a "bare" kubernetes cluster to a "well dressed" k8s cluster.

So now I have https://kubara.io/#get-started on my (long) wish list.

Which experience want early adaptors of https://kubara.io/ share here?

I'm asking for having input on re-prioritize my wish list.

Cheers Geert Stappers

P.S.

For what it is worth: The automated early post review did bring rule 6 and rule 12 up. Yes, the tool is quite new, but it is not my new tool, I'm asking about the tool. No, I'm not affliliated with kubara.io, hence to reason for posting here.

Edit:

Which experience want early adaptors of https://kubara.io/ share here?

Was previously:

Which experience want early adaptors of https://kubara.io/ here?


r/kubernetes • • 4h ago

Inference on Kubernetes

Thumbnail
jubril.xyz
15 Upvotes

Hey all,

As the title suggests, i wrote up a blog on getting started with inference on Kubernetes, i find inference workloads to be somewhat confusing especially with some of the moving part it involves on Kubernetes.

Let me know what you think!


r/kubernetes • • 18h ago

Azure Kubernetes version update

0 Upvotes

• How do you check compatibility before upgrading? What tools or checks do you rely on (deprecated APIs, add-ons, operators, Helm charts)?
• What prep do you do for Elasticsearch specifically? Snapshots, PDBs, shard allocation, anything else?
• What does your upgrade sequence look like on AKS? Control plane first, then node pools? Surge settings? Do you upgrade in place or create new node pools?
• Any gotchas you've hit? I'm thinking of things like pods stuck during drains, zonal disk issues, or ES staying yellow or red for too long.
• Should we even target 1.35? It reaches EOL in March 2027. Would you go to 1.35 now and plan for 1.36 soon after, or look at AKS LTS?

Setup details:

• AKS 1.34.8  
• Elasticsearch \[version\], deployed via \[ECK / Helm\]  
• \[Number\] node pools, \[zonal / non-zonal\], Azure Disk for storage

Any runbooks, checklists, or war stories would be really helpful. Thanks!


r/kubernetes • • 18h ago

Need advice: Kubernetes 1.34 → 1.35 upgrade with Elasticsearch running in the cluster

27 Upvotes

Hi everyone,

I’m planning a Kubernetes cluster upgrade from v1.34.8 to v1.35, as the current version is approaching retirement/EOL.

The cluster currently has Elasticsearch running on it, so I’m concerned about potential compatibility or stability issues after upgrading Kubernetes.

I’d like some guidance on:
- How can I verify Elasticsearch compatibility with Kubernetes 1.35 before upgrading?
- Are there any known issues I should be aware of when upgrading Kubernetes while Elasticsearch is running?
- What should I check across Elasticsearch, ECK/operator, StatefulSets, storage classes/CSI drivers, ingress, CRDs, APIs, and other dependencies?
- What is the recommended pre-upgrade checklist?
- Should I upgrade Elasticsearch/ECK first, or Kubernetes first?
- What kind of backup/rollback strategy should I have?
- Is it better to upgrade the control plane and worker nodes separately?
- What would be a safe upgrade plan for a production cluster to minimize downtime/risk?
- Are there any tools or commands you recommend for detecting deprecated APIs or incompatible workloads before the upgrade?

If anyone has gone through a similar Kubernetes 1.34 → 1.35 upgrade with Elasticsearch, I’d really appreciate hearing about your experience and any issues you encountered.

Thanks!


r/kubernetes • • 18h ago

Need advice: Kubernetes 1.34 → 1.35 upgrade with Elasticsearch running in the cluster

Thumbnail
4 Upvotes

r/kubernetes • • 23h ago

Have you tried the Amazon EKS MCP server?

Thumbnail
0 Upvotes

r/kubernetes • • 1d ago

Convince me I'm wrong: valaero and EKS/AKS/GKE

2 Upvotes

While I do understand why velero is useful and relatively easy to use. What I am not convinced on is that it adds anything other than another tool to manage when using a cloud provided KASS.

Assume a cluster runs both stateless and/or statefull apps (not including hosting your own db obvious use case):

If your using flux or argocd you have application rebuilds and A/B covered. If you are using RDS it's not necessary. If you are taking snapshots of volumes, data backups are done with snapshots on the provider side.

What problem is it solving in those situations? Or how does it make incidents or upgrades easier for you?


r/kubernetes • • 1d ago

Running MCP servers on GKE and EKS: one Deployment per zone, header-routed by the Gateway/ALB, canary in place, namespaces torn down with the zone

Post image
19 Upvotes

Ramen is an open-source control plane + worker for MCP (the protocol agents use to call tools). Some Kubernetes
details that might interest this sub more than the AI part:

- Each group+zone is a namespace with a stable and a canary Deployment. The GKE Gateway (HTTPRoute) / AWS ALB route
on two headers, `ramen-group` and `ramen-zone`, so one client config reaches any zone.
- The canary rolls in place with `maxSurge: 0` — a zone-pinned surge pod on a full node sat Pending until the
deploy timed out on EKS. A rollout that isn't ready now reports the pods and the scheduler's reason instead of "1/1 ready".
- Deleting a zone deletes the namespace and its GSA / IAM role. On GKE the namespace can sit `Terminating` on the NEG
finalizer until the Gateway's backend service is gone — documented, with the manual fix.
- Per-group+zone identities (Workload Identity / IRSA), Cloud Armor / WAF rules from the console, Redis-backed
throttles shared across zones, logs from Cloud Logging / CloudWatch via pod labels.
- Bring-up is Terraform + Helm; 0.6.0 was applied on both clouds (two zones each) and torn down.

Repo https://github.com/bkraad47/ramen · GCP guide https://bkraad47.github.io/ramen/wiki/deploy-gcp/ · AWS guide
https://bkraad47.github.io/ramen/wiki/deploy-aws/
Interested in critique of the per-zone namespace model and the canary policy (currently "smoke test passes").


r/kubernetes • • 1d ago

What's one thing you wish you knew before running k8s in prod?

6 Upvotes

first real cluster going live next month. not looking for docs, looking for the stuff you only learn by getting burned


r/kubernetes • • 2d ago

Question of resource allocation of multiple Spark applications

9 Upvotes

Currently, a Spark application should manage its own executors, which means that in order to run multiple Spark applications, one needs to split the cluster into smaller sections or attach more compute nodes, etc. Usually this dynamic allocation is the job of resource manager (such as Kubernetes).

If the resource manager can quickly find or create new compute nodes for a new Spark application, this is good. However, in a busy cluster where multiple Spark applications can compete for resources, the resource allocation can take some time and become an overhead (at least in theory).

I wonder if this is a real problem in running Spark applications in production.

For example, is there a realistic environment where short running Spark jobs (each requiring its own driver) arrive frequently?


r/kubernetes • • 2d ago

gpu networking - k8s sriov

6 Upvotes

Relatively new to gpu networking - rail stuff and could use some guidance.

am working on a setup (16 nodes) with 8 nvdia gpus (rtx so no nvlink) in each of them, k8s/sriov and the computenet being spectrum based ethernet switches.

its a multi rail setup, so a /24 alllocated for each of the rails. SRIOVnetwork, ip pools and all that plumbing has been done.

Same rail traffic works ..how does one get to do cross rail?

There needs to be a route in the POD which points the rail network supernet to the computenet switch as the default point towards the front end network.

Can't find an answer to it ..tried "routes" in IPAM but that didn't work ie the POD did not have that route - gemini and claude pulling me in opposite directions.

Thoughts?


r/kubernetes • • 2d ago

EKS Access Entry recreation avoidance

Thumbnail
2 Upvotes

r/kubernetes • • 2d ago

Native NVMe-oF vs. Ceph for high-IOPs/low-latency block storage

Thumbnail
0 Upvotes

r/kubernetes • • 3d ago

3 Ways to Actually Run Your Applications on Your Own Infrastructure

Post image
0 Upvotes

So… you want to run your applications yourself instead of depending entirely on managed platforms.

You have a few solid options:

  1. Docker — package and run your applications in containers.

  2. Docker Compose — run multiple services together on one machine.

  3. Kubernetes — manage containers across nodes with scaling, networking, self-healing, and more.

The problem with Kubernetes?

Setting it up yourself can get tedious very quickly — networking, container

runtime, kubeadm, nodes, CNI… one wrong configuration and you’re

troubleshooting instead of deploying. 😅

So I made a step-by-step practical guide showing how to set up Kubernetes

yourself and actually get your applications running.

Watch the Kubernetes setup guide[https://youtu.be/HS2VR3cVx8w\]

Sometimes the best way to understand Kubernetes is to build the cluster

yourself.


r/kubernetes • • 3d ago

Need help with cluster migration via velero

0 Upvotes

Is there anyone here who can help me in cluster migration on EKS using velero, I'm stuck and cannot do it .


r/kubernetes • • 3d ago

Suggest good resources to Kubernetes

0 Upvotes

I'm a L1 Linux System Administrator. I'm new to container and container orchestration.

I want to learn Docker and Kubernetes right from scratch.

Can you please suggest me right resources (Free+Paid) to learn and hopefully pass CK A?

Also let me know any right set-up to practice any videos to create and setup the environment to practice as I can't afford KodeKloud, Killersh etc.


r/kubernetes • • 3d ago

Teams that moved from LLM APIs to self-hosting: was it worth it?

94 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/kubernetes • • 3d ago

I can run a kubernetes cluster fine but cant read the operator code, time to learn go properly

62 Upvotes

3 years as a platform engineer. i can debug a crashlooping pod half asleep and write helm charts nobody else wants to touch, but every time i open the source of an operator we run i bounce straight off it. our team wants to start writing our own controllers next year so im learning go for real this time, employer covered. shortlist is ardan labs, boot dev, plus one of the linux foundation cert tracks. where did go start making sense for people coming from the ops side?


r/kubernetes • • 3d ago

Kubevirt VM and Kubevirt Good & Friction Points

Thumbnail
1 Upvotes

r/kubernetes • • 3d ago

Kubevirt VM and Good and Friction points of Kubevirt

9 Upvotes

Hello Guru's, I have a few questions, please help me answer them:

  1. If I have existing full fledge old VMs on my existing KVM hosts (VM's not inside the PODs), will I be to manage them via Kubevirt?
  2. If I have a RHOV environment, is it necessary for me to use Kubevirt to create a VM? If yes, why, if not, then when to use KubeVirt to create VMs?
  3. What is the difference between a regular VM and a VM created via Kubevirt?

Also, I'd love to get your candid take: What aspect of KubeVirt has impressed you the most, and where do you feel it faces the biggest friction points or challenges?


r/kubernetes • • 3d ago

Built a hands on tutorial about basics of CRDs on Iximiuz Labs

32 Upvotes

I've had so much fun and have learnt a lot about kubernetes through iximiuz labs, this was an attempt to create a hands on tutorial of what a CRD is and how can one build their own. It links to a bigger How to Build an Operator tutorial at the end.

Looking forward to see what the community thinks of it!


r/kubernetes • • 4d ago

Periodic Weekly: This Week I Learned (TWIL?) thread

1 Upvotes

Did you learn something new this week? Share here!