Phase II – Config Compilation Engine
Automate Triton Configs with Jinja2
Hand-crafting Triton’s config.pbtxt files for multiple models leads to broken setups. Instead, build a simple script to generate these files automatically from a master template.
Start by writing a Jinja2 template (templates/config.pbtxt.j2) that handles rules for dynamic batching, multi-instance GPU scaling, LLM streaming, and input/output tensor shapes.
[ model_catalog.yaml ]
│
▼
[ scripts/compile_configs.py ] ◄── [ templates/config.pbtxt.j2 ]
│
▼
[ model_repository/<model>/config.pbtxt ]
Next, define all your models in a central catalog (model_catalog.yaml). This file tracks setup details for every model, including qwen_1b (vLLM), bge_small_embedding (ONNX), and resnet50_vision (ONNX).
Use engine configuration files like model.json to pass settings directly to backend runtimes—for example, capping vLLM VRAM usage at 75% so it leaves memory for other tasks.
Run compile_configs.py using uv to build the configuration files. The script reads the catalog, creates model version folders (model_repository/<model>/1/), and writes out valid config.pbtxt files.