All Products
Search
Document Center

Container Service for Kubernetes:Support for inference frameworks

Last Updated:Sep 29, 2026

Gateway with Inference Extension supports multiple generative AI inference frameworks and delivers consistent capabilities across all of them, including canary releases, inference load balancing, and model name-based routing.

Supported inference frameworks

Inference framework Required version Notes
vLLM v0 >= v0.6.4 No additional configuration required.
vLLM v1 >= v0.8.0 No additional configuration required.
SGLang >= v0.3.6 Add the inference.networking.x-k8s.io/model-server-runtime: sglang annotation to your InferencePool.
Triton with a TensorRT-LLM backend >= 25.03 Add the inference.networking.x-k8s.io/model-server-runtime: trt-llm annotation to your InferencePool.

vLLM support

vLLM is the default backend inference framework. No additional configuration is required to enable intelligent routing and load balancing for vLLM-based inference services.

SGLang support

Add the inference.networking.x-k8s.io/model-server-runtime: sglang annotation to your InferencePool resource. No other resources require changes.

apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
  annotations:
    inference.networking.x-k8s.io/model-server-runtime: sglang
  name: deepseek-sglang-pool
spec:
  extensionRef:
    group: ""
    kind: Service
    name: deepseek-sglang-ext-proc
  selector:
    app: deepseek-r1-sglang
  targetPortNumber: 30000

TensorRT-LLM support

TensorRT-LLM is an open source engine from NVIDIA for optimizing LLM inference on NVIDIA GPUs. It builds TensorRT engines that run across one or more GPUs and supports Tensor Parallelism and Pipeline Parallelism. TensorRT-LLM integrates with Triton as a backend through the TensorRT-LLM Backend.

Add the inference.networking.x-k8s.io/model-server-runtime: trt-llm annotation to your InferencePool resource. No other resources require changes.

apiVersion: inference.networking.x-k8s.io/v1alpha2
kind: InferencePool
metadata:
  annotations:
    inference.networking.x-k8s.io/model-server-runtime: trt-llm
  name: qwen-trt-pool
spec:
  extensionRef:
    group: ""
    kind: Service
    name: trt-llm-ext-proc
  selector:
    app: qwen-trt-llm
  targetPortNumber: 8000