Phase III – Containerization & Local Deployment

Kubernetes Triton Inference Server vLLM License

Fix Container Inter-Process Communication (IPC) and Manage VRAM

Build a custom Docker image (triton-custom:24.08) that bundles Triton Server with PyTorch, vLLM dependencies, and ONNX Runtime backends.

For local deployments on K3s, write Kubernetes manifests in deploy/overlays/local-k3s/. Mount model folders directly using hostPath and launch Triton with explicit model control flags (--load-model) so it only loads specified models at startup.

Expose port 30800 for HTTP requests, 30801 for gRPC calls, and 30802 for Prometheus metrics using a NodePort Service.

Client Request
      │
      ▼
NodePort Service (30800 / 30801)
      │
      ▼
Triton Pod (Explicit Loading Mode)
   ├── Shared Memory RAM Mount (/dev/shm)
   ├── ONNX: Embeddings & Vision
   └── vLLM: Qwen LLM Streaming

After deploying, run kubectl logs to verify startup steps and check POST /v2/repository/index to see loaded models. Test the running server by sending sample requests with curl for text embeddings, image classification, and streaming LLM responses.

More