Gateway with Inference Extension supports multiple generative AI inference frameworks and delivers consistent capabilities across all of them, including canary releases, inference load balancing, and model name-based routing.
Supported inference frameworks
| Inference framework | Required version | Notes |
|---|---|---|
| vLLM v0 | >= v0.6.4 | No additional configuration required. |
| vLLM v1 | >= v0.8.0 | No additional configuration required. |
| SGLang | >= v0.3.6 | Add the inference.networking.x-k8s.io/model-server-runtime: sglang annotation to your InferencePool. |
| Triton with a TensorRT-LLM backend | >= 25.03 | Add the inference.networking.x-k8s.io/model-server-runtime: trt-llm annotation to your InferencePool. |
vLLM support
vLLM is the default backend inference framework. No additional configuration is required to enable intelligent routing and load balancing for vLLM-based inference services.
SGLang support
Add the inference.networking.x-k8s.io/model-server-runtime: sglang annotation to your InferencePool resource. No other resources require changes.
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
annotations:
inference.networking.x-k8s.io/model-server-runtime: sglang
name: deepseek-sglang-pool
spec:
extensionRef:
group: ""
kind: Service
name: deepseek-sglang-ext-proc
selector:
app: deepseek-r1-sglang
targetPortNumber: 30000
TensorRT-LLM support
TensorRT-LLM is an open source engine from NVIDIA for optimizing LLM inference on NVIDIA GPUs. It builds TensorRT engines that run across one or more GPUs and supports Tensor Parallelism and Pipeline Parallelism. TensorRT-LLM integrates with Triton as a backend through the TensorRT-LLM Backend.
Add the inference.networking.x-k8s.io/model-server-runtime: trt-llm annotation to your InferencePool resource. No other resources require changes.
apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
annotations:
inference.networking.x-k8s.io/model-server-runtime: trt-llm
name: qwen-trt-pool
spec:
extensionRef:
group: ""
kind: Service
name: trt-llm-ext-proc
selector:
app: qwen-trt-llm
targetPortNumber: 8000