Cloud-Agnostic Multi-Modal Model Serving Engine
This guide walks through setting up a local multi-modal model serving stack on single-node K3s and migrating it to AWS EKS. It replaces single-user CLI wrappers with Triton Inference Server, vLLM, and ONNX Runtime to support dynamic batching, shared GPU memory management, and GitOps deployments.
The general consensus online is that running local models has become trivial. That is certainly true if the goal is running a single-user cli wrapper like ollama run in a terminal window. However, that abstraction falls apart the moment one need dynamic batching, heterogenous multi-model execution, optimized GPU memory lifecycle management, or the ability to manage host CUDA drivers predictably without breaking local virtual environments.
Rather than cluttering the host os with ad-hoc python scripts or standard homelab setups, I began building a local k3s environment leveraging NVIDIA Triton Inference Server with vLLM and ONNX Runtime backends. The goal is to establish a local edge execution environment that mirrors the exact containerized patterns, declarative configurations, and model registry topologies required for production deployments on Amazon EKS. For detailed technical steps, check out the README.
+---------------------------------------------------------------------------------------------------+
| LAYER 1: CONFIG COMPILATION ENGINE |
| [ model_catalog.yaml ] + [ templates/config.pbtxt.j2 ] |
| │ |
| ▼ (uv run python scripts/compile_configs.py) |
| [ Compiled Model Repository Artifacts ] ──> model_repository/<model>/config.pbtxt |
+---------------------------------------------------------------------------------------------------+
│
│ Bakes directly into Docker image / Mounts via HostPath
▼
+---------------------------------------------------------------------------------------------------+
| LAYER 2: LOCAL K3s CLUSTER - SINGLE NODE (NVIDIA RTX 3070 Ti) |
| [ NodePort Ingress Services: HTTP (30800) | gRPC (30801) | Metrics (30802) ] |
| │ |
| ▼ |
| [ Triton Inference Server Pod (Namespace: triton) ] |
| ├── Image: triton-custom:24.08 |
| ├── Container Runtime: Containerd (runtimeClassName: nvidia) |
| ├── Memory IPC: Shared Memory /dev/shm (Mounted via RAM emptyDir) |
| ├── Flags: --model-control-mode=explicit --load-model=simple ... |
| │ |
| ├── ONNX Runtime Backend ──> simple (Dynamic Batching, Instances: 2) |
| ├── ONNX Runtime Backend ──> resnet50_vision (Dynamic Batching, Instances: 2) |
| └── vLLM C++ Backend ──> qwen_1b (Decoupled Streaming, model.json 55% VRAM) |
| |
| [ Shared VRAM Hardware Pool (8 GB) ] ──> Dynamic Reclaim via POST /v2/repository/.../unload |
+---------------------------------------------------------------------------------------------------+
│
│ Git Push / Helm Chart & ArgoCD Sync
▼
+---------------------------------------------------------------------------------------------------+
| LAYER 3: GITOPS & CLOUD MIGRATION TARGET (AWS EKS) |
| [ Git Repository (krishpn/cicdtest) ] |
| ├── helm/triton-inference-server/ (Core Helm Chart & Templates) |
| └── gitops/triton-app.yaml (ArgoCD Declarative Application Spec) |
| │ |
| ▼ |
| [ ArgoCD Continuous Delivery Engine ] ──> [ AWS EKS Production Infrastructure ] |
+---------------------------------------------------------------------------------------------------+