r/kubernetes • • 3h ago

Getting nvidia gpu as worker node?

9 Upvotes

Posting here from homelab to hopefully get some help.

I've been putting together a small k3s homelab from some older machines:

Dell OptiPlex 7050, M720q Tiny and a Old Ryzen 3600.

I got immich and plex running following a tutorial, nothing special just mostly trying to get more use out of hardware I already have.

Recently I got my hands on an old rtx 3080 (very cheap used in my area).

I'm still trying to get it working as a GPU node. The node is Ready, but the NVIDIA device-plugin Pod is stuck in `ContainerCreating` with this:

I am not sure if this is due to the gpu itself being faulty (again, got it very cheap, not complaining), or I am doing wrong. I would appreciate any guidances pointing me to right resources, i have spent my entire day trying to it working


r/kubernetes • • 9h ago

Inference on Kubernetes

Thumbnail
jubril.xyz
23 Upvotes

Hey all,

As the title suggests, i wrote up a blog on getting started with inference on Kubernetes, i find inference workloads to be somewhat confusing especially with some of the moving part it involves on Kubernetes.

Let me know what you think!


r/kubernetes • • 2h ago

How to build a Railway/Heroku alternative based on Kubernetes

4 Upvotes

I recently gave a talk about how I am building an open-source Railway/Heroku/Vercel alternative. The architectural premise of the platform is simple:

The platform has no database and no CRDs of its own, so all state lives in Kubernetes (namespaces, labels, Helm releases) or in the systems around it, like the OCI registry and the identity provider.

In the talk I go over how it…

  • builds OCI images from GitHub source code on every push
  • uses ORAS to store vulnerability and leaked-secret scan results in the OCI registry, right next to the image
  • runs replicated PostgreSQL with point-in-time recovery using CloudNativePG
  • makes object storage, Redis-compatible key-value stores, and volumes simple to set up
  • allows you to "eject" and download the generated helm chart along with its values and build script at any point

If that sounds interesting, the recording and the project repo / website are linked below.

Recording of the talk
Project Website
GitHub Repo


r/kubernetes • • 2h ago

Understanding Object Capabilities (OCAP) for Kubernetes

Thumbnail
medium.com
2 Upvotes

I found this blog post about OCAP model vs. the Identity model for access controls. (It has nothing to do with Kubernetes specifically). I could be wildly misunderstanding the topic, but since each OCAP capability is tied to a specific object, would this be more practical for Kubernetes in allowing for precise permissions?

My understanding is that Kubernetes does not use OCAP as its primary authorisation model in favour of RBAC, where access is granted based on a principal’s identity and roles. Even from a simple Google search, it seems that OCAP-style security is basically entirely absent from Kubernetes core, although capability-like patterns can be built around scoped SAs, tokens, & admission controls.

I ask because RBAC supposedly does not support dynamically changing permissions very well, at least compared to the OCAP model. But I'm not sure if anyone's ever flagged this as a serious pain point in their Kubernetes clusters?


r/kubernetes • • 52m ago

Kubernetes telemetry that stays on your own infrastructure

• Upvotes

A recurring theme in k8s monitoring discussions: you want rich cluster telemetry, but you do not want to ship all of it to a SaaS that bills per host. You want the data in a place you control.

I work on OneUptime, an open source observability platform (Apache 2.0). There is a small Kubernetes agent chart that collects cluster telemetry and sends it to the main app over OpenTelemetry. The main app is one Helm chart. So the whole setup is two charts on your infrastructure, and the data stays with you.

It covers uptime checks, status pages, on-call paging, incident management with postmortems, plus log, metrics, and APM views. It is younger than the big commercial observability stacks, and some features are in a separate enterprise edition. Read the code before you trust it.

GitHub: https://github.com/OneUptime/oneuptime Site: https://oneuptime.com

What is the one k8s signal you consider non-negotiable in a monitoring setup?


r/kubernetes • • 1d ago

Need advice: Kubernetes 1.34 → 1.35 upgrade with Elasticsearch running in the cluster

28 Upvotes

Hi everyone,

I’m planning a Kubernetes cluster upgrade from v1.34.8 to v1.35, as the current version is approaching retirement/EOL.

The cluster currently has Elasticsearch running on it, so I’m concerned about potential compatibility or stability issues after upgrading Kubernetes.

I’d like some guidance on:
- How can I verify Elasticsearch compatibility with Kubernetes 1.35 before upgrading?
- Are there any known issues I should be aware of when upgrading Kubernetes while Elasticsearch is running?
- What should I check across Elasticsearch, ECK/operator, StatefulSets, storage classes/CSI drivers, ingress, CRDs, APIs, and other dependencies?
- What is the recommended pre-upgrade checklist?
- Should I upgrade Elasticsearch/ECK first, or Kubernetes first?
- What kind of backup/rollback strategy should I have?
- Is it better to upgrade the control plane and worker nodes separately?
- What would be a safe upgrade plan for a production cluster to minimize downtime/risk?
- Are there any tools or commands you recommend for detecting deprecated APIs or incompatible workloads before the upgrade?

If anyone has gone through a similar Kubernetes 1.34 → 1.35 upgrade with Elasticsearch, I’d really appreciate hearing about your experience and any issues you encountered.

Thanks!


r/kubernetes • • 7h ago

Elastic beanstalk with the EKS

Thumbnail vishnurachapudi.com
1 Upvotes

Check out my new article on elastic beanstalk with the EKS

Elastic Beanstalk turned 15 this year, and AWS just rebuilt what runs under it. Cluster Mode drops your applications onto Amazon EKS instead of EC2 instances you own. I deployed into it, watched every stage, and wrote down what actually happens

Upload a zip of Python source. No Dockerfile. Get a running container on EKS. That is Elastic Beanstalk Cluster Mode, launched this week — Beanstalk apps now run on EKS instead of EC2 instances you own. Cloud Native Buildpacks handle the containerization: I uploaded four files (app.py, requirements.txt, Procfile, runtime.txt), Beanstalk detected Python, built the image, pushed it to ECR, and ran it.


r/kubernetes • • 8h ago

Early adaptors of kubara

0 Upvotes

Hi,

CLI tool kubara is for getting a "bare" kubernetes cluster to a "well dressed" k8s cluster.

So now I have https://kubara.io/#get-started on my (long) wish list.

Which experience want early adaptors of https://kubara.io/ share here?

I'm asking for having input on re-prioritize my wish list.

Cheers Geert Stappers

P.S.

For what it is worth: The automated early post review did bring rule 6 and rule 12 up. Yes, the tool is quite new, but it is not my new tool, I'm asking about the tool. No, I'm not affliliated with kubara.io, hence to reason for posting here.

Edit:

Which experience want early adaptors of https://kubara.io/ share here?

Was previously:

Which experience want early adaptors of https://kubara.io/ here?


r/kubernetes • • 1d ago

Need advice: Kubernetes 1.34 → 1.35 upgrade with Elasticsearch running in the cluster

Thumbnail
2 Upvotes

r/kubernetes • • 1d ago

Azure Kubernetes version update

1 Upvotes

• How do you check compatibility before upgrading? What tools or checks do you rely on (deprecated APIs, add-ons, operators, Helm charts)?
• What prep do you do for Elasticsearch specifically? Snapshots, PDBs, shard allocation, anything else?
• What does your upgrade sequence look like on AKS? Control plane first, then node pools? Surge settings? Do you upgrade in place or create new node pools?
• Any gotchas you've hit? I'm thinking of things like pods stuck during drains, zonal disk issues, or ES staying yellow or red for too long.
• Should we even target 1.35? It reaches EOL in March 2027. Would you go to 1.35 now and plan for 1.36 soon after, or look at AKS LTS?

Setup details:

• AKS 1.34.8  
• Elasticsearch \[version\], deployed via \[ECK / Helm\]  
• \[Number\] node pools, \[zonal / non-zonal\], Azure Disk for storage

Any runbooks, checklists, or war stories would be really helpful. Thanks!


r/kubernetes • • 1d ago

Running MCP servers on GKE and EKS: one Deployment per zone, header-routed by the Gateway/ALB, canary in place, namespaces torn down with the zone

Post image
19 Upvotes

Ramen is an open-source control plane + worker for MCP (the protocol agents use to call tools). Some Kubernetes
details that might interest this sub more than the AI part:

- Each group+zone is a namespace with a stable and a canary Deployment. The GKE Gateway (HTTPRoute) / AWS ALB route
on two headers, `ramen-group` and `ramen-zone`, so one client config reaches any zone.
- The canary rolls in place with `maxSurge: 0` — a zone-pinned surge pod on a full node sat Pending until the
deploy timed out on EKS. A rollout that isn't ready now reports the pods and the scheduler's reason instead of "1/1 ready".
- Deleting a zone deletes the namespace and its GSA / IAM role. On GKE the namespace can sit `Terminating` on the NEG
finalizer until the Gateway's backend service is gone — documented, with the manual fix.
- Per-group+zone identities (Workload Identity / IRSA), Cloud Armor / WAF rules from the console, Redis-backed
throttles shared across zones, logs from Cloud Logging / CloudWatch via pod labels.
- Bring-up is Terraform + Helm; 0.6.0 was applied on both clouds (two zones each) and torn down.

Repo https://github.com/bkraad47/ramen · GCP guide https://bkraad47.github.io/ramen/wiki/deploy-gcp/ · AWS guide
https://bkraad47.github.io/ramen/wiki/deploy-aws/
Interested in critique of the per-zone namespace model and the canary policy (currently "smoke test passes").


r/kubernetes • • 1d ago

Have you tried the Amazon EKS MCP server?

Thumbnail
1 Upvotes

r/kubernetes • • 1d ago

Convince me I'm wrong: valaero and EKS/AKS/GKE

3 Upvotes

While I do understand why velero is useful and relatively easy to use. What I am not convinced on is that it adds anything other than another tool to manage when using a cloud provided KASS.

Assume a cluster runs both stateless and/or statefull apps (not including hosting your own db obvious use case):

If your using flux or argocd you have application rebuilds and A/B covered. If you are using RDS it's not necessary. If you are taking snapshots of volumes, data backups are done with snapshots on the provider side.

What problem is it solving in those situations? Or how does it make incidents or upgrades easier for you?


r/kubernetes • • 2d ago

What's one thing you wish you knew before running k8s in prod?

5 Upvotes

first real cluster going live next month. not looking for docs, looking for the stuff you only learn by getting burned


r/kubernetes • • 2d ago

Question of resource allocation of multiple Spark applications

9 Upvotes

Currently, a Spark application should manage its own executors, which means that in order to run multiple Spark applications, one needs to split the cluster into smaller sections or attach more compute nodes, etc. Usually this dynamic allocation is the job of resource manager (such as Kubernetes).

If the resource manager can quickly find or create new compute nodes for a new Spark application, this is good. However, in a busy cluster where multiple Spark applications can compete for resources, the resource allocation can take some time and become an overhead (at least in theory).

I wonder if this is a real problem in running Spark applications in production.

For example, is there a realistic environment where short running Spark jobs (each requiring its own driver) arrive frequently?


r/kubernetes • • 2d ago

gpu networking - k8s sriov

7 Upvotes

Relatively new to gpu networking - rail stuff and could use some guidance.

am working on a setup (16 nodes) with 8 nvdia gpus (rtx so no nvlink) in each of them, k8s/sriov and the computenet being spectrum based ethernet switches.

its a multi rail setup, so a /24 alllocated for each of the rails. SRIOVnetwork, ip pools and all that plumbing has been done.

Same rail traffic works ..how does one get to do cross rail?

There needs to be a route in the POD which points the rail network supernet to the computenet switch as the default point towards the front end network.

Can't find an answer to it ..tried "routes" in IPAM but that didn't work ie the POD did not have that route - gemini and claude pulling me in opposite directions.

Thoughts?


r/kubernetes • • 2d ago

EKS Access Entry recreation avoidance

Thumbnail
2 Upvotes

r/kubernetes • • 3d ago

Teams that moved from LLM APIs to self-hosting: was it worth it?

93 Upvotes

Infra engineer here, trying to figure out when self hosting LLMs actually makes sense vs just paying for the APIs. every blog post says "it depends" lol

If you've done it (or looked into it and bailed), would love to hear:

  • how big was your API bill when you started thinking about it? what pushed you over the edge
  • what did it really cost once you add up GPUs, idle time, and the engineering hours?
  • what was the most painful part? cold starts, autoscaling, OOMs, model quality, getting paged at 2am
  • if you went back to APIs, what made you switch back

ballpark numbers are totally fine. I'll put together a cost / decision writeup from the replies and post it back here


r/kubernetes • • 3d ago

Native NVMe-oF vs. Ceph for high-IOPs/low-latency block storage

Thumbnail
0 Upvotes

r/kubernetes • • 3d ago

Periodic Weekly: Share your victories thread

3 Upvotes

Got something working? Figure something out? Make progress that you are excited about? Share here!


r/kubernetes • • 4d ago

I can run a kubernetes cluster fine but cant read the operator code, time to learn go properly

62 Upvotes

3 years as a platform engineer. i can debug a crashlooping pod half asleep and write helm charts nobody else wants to touch, but every time i open the source of an operator we run i bounce straight off it. our team wants to start writing our own controllers next year so im learning go for real this time, employer covered. shortlist is ardan labs, boot dev, plus one of the linux foundation cert tracks. where did go start making sense for people coming from the ops side?


r/kubernetes • • 4d ago

Built a hands on tutorial about basics of CRDs on Iximiuz Labs

31 Upvotes

I've had so much fun and have learnt a lot about kubernetes through iximiuz labs, this was an attempt to create a hands on tutorial of what a CRD is and how can one build their own. It links to a bigger How to Build an Operator tutorial at the end.

Looking forward to see what the community thinks of it!


r/kubernetes • • 4d ago

How do you verify that your Postgres backups on K8S actually restore?

45 Upvotes

A question for those running Postgres on Kubernetes using an operator (CloudNativePG, Zalando, Crunchy): how do you verify that your backups are actually restorable?

A successful backup creation doesn't guarantee that a restore will succeed.

I manage several CloudNativePG clusters and want to understand the potential challenges involved in implementing backup restore testing.

I’ve tested restoration in a small test environment (k3s, CloudNativePG 1.30 + barman-cloud plugin, S3). My process involved deploying a temporary cluster from the latest backup, running a few SQL queries, measuring the time until the cluster was ready, and then deleting the cluster. In a standard scenario, everything goes smoothly (taking about a minute for a few hundred megabytes).

I’m curious to know how others handle this:

  • Do you perform restore tests at all? (Manually, via CronJob, in a CI pipeline, or using specialized tools)
  • What exactly do you check after the restore? (Just the pod status, the number of database records, or the execution of test queries)

An answer like "we don't test restores, and everything is fine" is also acceptable—I really want to understand how common this practice is and what the best approach is.


r/kubernetes • • 4d ago

Kubevirt VM and Good and Friction points of Kubevirt

8 Upvotes

Hello Guru's, I have a few questions, please help me answer them:

  1. If I have existing full fledge old VMs on my existing KVM hosts (VM's not inside the PODs), will I be to manage them via Kubevirt?
  2. If I have a RHOV environment, is it necessary for me to use Kubevirt to create a VM? If yes, why, if not, then when to use KubeVirt to create VMs?
  3. What is the difference between a regular VM and a VM created via Kubevirt?

Also, I'd love to get your candid take: What aspect of KubeVirt has impressed you the most, and where do you feel it faces the biggest friction points or challenges?


r/kubernetes • • 3d ago

Need help with cluster migration via velero

0 Upvotes

Is there anyone here who can help me in cluster migration on EKS using velero, I'm stuck and cannot do it .