Phase I- Host Prerequisites & K3s Runtime Patching
Configure GPU Access on Bare-Metal K3s
First, set up your local workspace by creating directories for code, deployment files, scripts, templates, and model files (apps/, deploy/, model_repository/, templates/, and scripts/). Initialize Git and add model binaries like .onnx, .engine, and .safetensors to your .gitignore so large weights do not bloat repository. Use .gitkeep files to track empty folders.
Next, give your containers access to the host GPU. Install the NVIDIA container toolkit and its repository keys, then run nvidia-ctk to verify the drivers expose your hardware to container runtimes.
Host Drivers & GPU Hardware
│
▼
NVIDIA Container Toolkit (nvidia-ctk)
│
▼
Containerd Runtime (/usr/bin/nvidia-container-runtime)
│
▼
K3s Pods (runtimeClassName: nvidia)
By default, K3s puts networking files and sockets in non-standard folders, which breaks GPU plugins and network setups when the host reboots. Fix this by pointing Flannel’s CNI config paths to /etc/cni/net.d and its binaries to /opt/cni/bin.
Then, link Kubelet’s device plugin socket to /var/lib/kubelet/device-plugins. Finally, edit /var/lib/rancher/k3s/agent/etc/containerd/config.toml.tmpl so Containerd uses /usr/bin/nvidia-container-runtime as its default runtime.
Once K3s restarts, deploy the NVIDIA Kubernetes Device Plugin DaemonSet and run kubectl get nodes to confirm the cluster shows available nvidia.com/gpu resources.