Phase III – Containerization & Local Deployment
Fix Container Inter-Process Communication (IPC) and Manage VRAM
Build a custom Docker image (triton-custom:24.08) that bundles Triton Server with PyTorch, vLLM dependencies, and ONNX Runtime backends.
For local deployments on K3s, write Kubernetes manifests in deploy/overlays/local-k3s/. Mount model folders directly using hostPath and launch Triton with explicit model control flags (--load-model) so it only loads specified models at startup.
Expose port 30800 for HTTP requests, 30801 for gRPC calls, and 30802 for Prometheus metrics using a NodePort Service.
Client Request
│
▼
NodePort Service (30800 / 30801)
│
▼
Triton Pod (Explicit Loading Mode)
├── Shared Memory RAM Mount (/dev/shm)
├── ONNX: Embeddings & Vision
└── vLLM: Qwen LLM Streaming
After deploying, run kubectl logs to verify startup steps and check POST /v2/repository/index to see loaded models. Test the running server by sending sample requests with curl for text embeddings, image classification, and streaming LLM responses.