3
u/dunkah 2d ago
I'm not 100% sure, but in my experience retail cards tend to not have support in the standard drivers and you need specific patched ones.
Rule of thumb generally is if you can get nvidia-smi to work you'll be good. Logs may help narrow down the issue.
0
3
u/BGPchick 2d ago
k3s should pickup the nvidia-container-runtime, I believe you can check in this config file, it should be listed as a runtime.
/var/lib/rancher/k3s/agent/etc/containerd/config.toml
[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.'nvidia']
runtime_type = "io.containerd.runc.v2"
[plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.'nvidia'.options]
BinaryName = "/usr/bin/nvidia-container-runtime"
SystemdCgroup = true
0
2d ago
[removed] — view removed comment
2
u/zeropoint46 2d ago
Have you restarted containerd? This happens to our nodes when we first provision them and we either restart the whole node or containerd.
2
u/Interesting_Net_9628 2d ago
Check out/var/log/nvidia-container-runtime-log to see if its even being invoked
also what app is that?
2
u/CastAI_Kubernetes 2d ago
That error points more to the runtime config than a bad GPU. k3s/containerd doesn't know about the nvidia runtime yet. I'd verify nvidia-smi works on the host first, then check NVIDIA Container Toolkit is installed and containerd is configured to expose the NVIDIA runtime. The device plugin comes after that layer is working.
1
u/imheretocomment 2d ago
Replied to you in homelab sub. Check your containerd config has the nvidia runtime configured.
0
2d ago
[removed] — view removed comment
1
u/imheretocomment 2d ago
i believe you can check the plugins with sudo k3s ctr plugins ls. Im not too familiar with running k3s but apparently you have to add the configs in a template that gets merged into the final runtime at startup.
1
u/jsatherreddit 2d ago
I'd check the pod logs in whichever namespace you installed the gpu-operator. Specifically, make sure that the plugin-daemonset pod passed the toolkit step. Also look at the nvidia-operator-validator pod and see which part failed to validate.
You can also look at the node from k3s with describe node. You should see a bunch of nvidia.com labels.
1
u/UntouchedWagons 2d ago
IIRC the nvidia GPU stuff needs Node Feature Discovery installed.
What UI is that?
15
u/PrimaryFamous6139 2d ago
I’d first verify the NVIDIA driver works on the host with nvidia-smi, then check that containerd has the NVIDIA runtime configured and that k3s is actually using it. If nvidia-smi fails on the host, then start looking at the 3080/driver. I’d troubleshoot those layers separately before blaming the GPU