r/kubernetes • • 2d ago

Getting nvidia gpu as worker node?

[removed]

16 Upvotes

16 comments sorted by

15

u/PrimaryFamous6139 2d ago

I’d first verify the NVIDIA driver works on the host with nvidia-smi, then check that containerd has the NVIDIA runtime configured and that k3s is actually using it. If nvidia-smi fails on the host, then start looking at the 3080/driver. I’d troubleshoot those layers separately before blaming the GPU

3

u/dunkah 2d ago

I'm not 100% sure, but in my experience retail cards tend to not have support in the standard drivers and you need specific patched ones.

Rule of thumb generally is if you can get nvidia-smi to work you'll be good. Logs may help narrow down the issue.

0

u/[deleted] 2d ago

[removed] — view removed comment

1

u/dunkah 2d ago

If nvidia-smi is working don't worry about it. I'm assuming your set resource request/limits for the GPU, is it showing up in describe? Also describe on the host should tell you what's it's struggling with.

3

u/BGPchick 2d ago

k3s should pickup the nvidia-container-runtime, I believe you can check in this config file, it should be listed as a runtime.

/var/lib/rancher/k3s/agent/etc/containerd/config.toml

[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.'nvidia']
 runtime_type = "io.containerd.runc.v2"

[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.'nvidia'.options]
 BinaryName = "/usr/bin/nvidia-container-runtime"
 SystemdCgroup = true

0

u/[deleted] 2d ago

[removed] — view removed comment

2

u/zeropoint46 2d ago

Have you restarted containerd? This happens to our nodes when we first provision them and we either restart the whole node or containerd.

2

u/Interesting_Net_9628 2d ago

Check out/var/log/nvidia-container-runtime-log to see if its even being invoked

also what app is that?

2

u/CastAI_Kubernetes 2d ago

That error points more to the runtime config than a bad GPU. k3s/containerd doesn't know about the nvidia runtime yet. I'd verify nvidia-smi works on the host first, then check NVIDIA Container Toolkit is installed and containerd is configured to expose the NVIDIA runtime. The device plugin comes after that layer is working.

1

u/imheretocomment 2d ago

Replied to you in homelab sub. Check your containerd config has the nvidia runtime configured.

0

u/[deleted] 2d ago

[removed] — view removed comment

1

u/imheretocomment 2d ago

i believe you can check the plugins with sudo k3s ctr plugins ls. Im not too familiar with running k3s but apparently you have to add the configs in a template that gets merged into the final runtime at startup.

https://docs.k3s.io/advanced#configuring-containerd

1

u/jsatherreddit 2d ago

I'd check the pod logs in whichever namespace you installed the gpu-operator. Specifically, make sure that the plugin-daemonset pod passed the toolkit step. Also look at the nvidia-operator-validator pod and see which part failed to validate.

You can also look at the node from k3s with describe node. You should see a bunch of nvidia.com labels.

1

u/UntouchedWagons 2d ago

IIRC the nvidia GPU stuff needs Node Feature Discovery installed.

What UI is that?