This page covers what's new in inference-nv-pytorch version 26.01.
What's new
Two CUDA variants. This release ships images for CUDA 12.8 (amd64 only) and CUDA 13.0 (amd64 and aarch64).
Upgraded inference frameworks. Both CUDA variants upgrade vLLM to v0.14.0 and SGLang to v0.5.7.
Updated DeepGPU components (CUDA 12.8 only). deepgpu-comfyui is upgraded to 1.4.1 and deepgpu-torch to 0.1.18+torch2.9.0cu128.
Bug fixes. None.
Image tags and system components
All images use the inference-nv-pytorch image name and target large model inference workloads with the PyTorch framework.
CUDA 12.8 images (amd64)
NVIDIA Driver release >= 570 required.
| Tag | Inference framework | System components |
|---|---|---|
26.01-vllm0.14.0-pytorch2.9-cu128-20260121-serverless | vLLM 0.14.0 | Ubuntu 24.04, Python 3.12, Torch 2.9.1, CUDA 12.8, diffusers 0.36.0, deepgpu-comfyui 1.4.1, deepgpu-torch 0.1.18+torch2.9.0cu128, flash_attn 2.8.3, flashinfer-python 0.5.3, imageio 2.37.2, imageio-ffmpeg 0.6.0, ray 2.53.0, transformers 4.57.6, triton 3.5.1, torchaudio 2.9.1, torchvision 0.24.1, vllm 0.14.0, xfuser 0.4.5, xgrammar 0.1.27, ljperf 0.1.0+477686c5 |
26.01-sglang0.5.7-pytorch2.9-cu128-20260113-serverless | SGLang 0.5.7 | Ubuntu 24.04, Python 3.12, Torch 2.9.1+cu128, CUDA 12.8, torchaudio 2.9.1+128, torchvision 0.24.1+128, diffusers 0.36.0, decord 0.6.0, decord2 3.0.0, deepgpu-comfyui 1.4.1, deepgpu-torch 0.1.18+torch2.9.0cu128, flash_attn 2.8.3, flash_mla 1.0.0+1408756, flashinfer-python 0.5.3, imageio 2.37.2, imageio-ffmpeg 0.6.0, ray 2.53.0, transformers 4.57.1, sgl-kernel 0.3.20, sglang 0.5.7, xgrammar 0.1.27, triton 3.5.1, torchao 0.9.0, xfuser 0.4.5, ljperf 0.1.0+477686c5 |
CUDA 13.0 images (amd64 and aarch64)
NVIDIA Driver release >= 580 required.
| Tag | Architecture | Inference framework | System components |
|---|---|---|---|
26.01-vllm0.14.0-pytorch2.9-cu130-20260123-serverless | amd64 | vLLM 0.14.0 | Ubuntu 24.04, Python 3.12, Torch 2.9.1+cu130, CUDA 13.0.2, diffusers 0.36.0, flash_attn 2.8.3, flashinfer-python 0.5.3, imageio 2.37.2, imageio-ffmpeg 0.6.0, ray 2.53.1, transformers 4.57.6, triton 3.5.0, torchaudio 2.9.1+cu130, torchvision 0.24.1+cu130, vllm 0.14.0, xfuser 0.4.5, xgrammar 0.1.27, ljperf 0.1.0+d0e4a408 |
26.01-vllm0.14.0-pytorch2.9-cu130-20260123-serverless | aarch64 | vLLM 0.14.0 | Ubuntu 24.04, Python 3.12, Torch 2.9.1+cu130, CUDA 13.0.2, diffusers 0.36.0, flash_attn 2.8.3, flashinfer-python 0.5.3, transformers 4.57.6, ray 2.53.0, vllm 0.14.0, triton 3.5.1, torchaudio 2.9.1+cu130, torchvision 2.9.1+cu130, xfuser 0.4.5, xgrammar 0.1.27, ljperf 0.1.0+477686c5 |
26.01-sglang0.5.7-pytorch2.9-cu130-20260113-serverless | amd64 | SGLang 0.5.7 | Ubuntu 24.04, Python 3.12, Torch 2.9.1+cu130, CUDA 13.0.2, diffusers 0.36.0, decord 0.6.0, decord2 3.0.0, flash_attn 2.8.3, flashinfer-python 0.5.3, imageio 2.37.2, imageio-ffmpeg 0.6.0, ray 2.53.0, transformers 4.57.1, sgl-kernel 0.3.20, sglang 0.5.7, xgrammar 0.1.27, triton 3.5.1, torchao 0.9.0, torchaudio 2.9.1, torchvision 0.24.1+cu130, xfuser 0.4.5, ljperf 0.1.0+d0e4a408 |
26.01-sglang0.5.7-pytorch2.9-cu130-20260113-serverless | aarch64 | SGLang 0.5.7 | Ubuntu 24.04, Python 3.12, Torch 2.9.1+cu130, CUDA 13.0.2, diffusers 0.36.0, decord2 3.0.0, flash_attn 2.8.3, flashinfer-python 0.5.3, imageio 2.37.2, imageio-ffmpeg 0.6.0, transformers 4.57.1, sgl-kernel 0.3.20, sglang 0.5.7, xgrammar 0.1.27, triton 3.5.1, torchao 0.9.0, torchaudio 2.9.1, torchvision 0.24.1, xfuser 0.4.5 |
Image addresses
Public network
CUDA 12.8
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-vllm0.14.0-pytorch2.9-cu128-20260121-serverless
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-sglang0.5.7-pytorch2.9-cu128-20260113-serverlessCUDA 13.0
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-vllm0.14.0-pytorch2.9-cu130-20260123-serverless
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-sglang0.5.7-pytorch2.9-cu130-20260113-serverlessVPC
To pull images from within a VPC, replace the public registry hostname with the VPC endpoint for your region:
Public:
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/{image:tag}VPC:
acs-registry-vpc.{region-id}.cr.aliyuncs.com/egslingjun/{image:tag}
Replace {region-id} with an ACS-supported region such as cn-beijing or cn-wulanchabu, and {image:tag} with the target image name and tag.
These images are compatible with ACS and Lingjun multi-tenant deployments. They are not supported in Lingjun single-tenant scenarios.
Quick start
The following example pulls the inference-nv-pytorch image and runs a vLLM inference test with the Qwen2.5-7B-Instruct model.
To use this image in ACS, select it from the Artifact Center page in the Workloads interface, or specify the image reference in a YAML file. For deployment guides, see:
Pull the container image.
docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]Download the Qwen2.5-7B-Instruct model from ModelScope.
pip install modelscope cd /mnt modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-InstructStart the container.
docker run -d -t --network=host --privileged --init --ipc=host \ --ulimit memlock=-1 --ulimit stack=67108864 \ -v /mnt/:/mnt/ \ egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]Start the vLLM server and test inference.
Start the server.
python3 -m vllm.entrypoints.openai.api_server \ --model /mnt/Qwen2.5-7B-Instruct \ --trust-remote-code --disable-custom-all-reduce \ --tensor-parallel-size 1Test from the client side.
curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "/mnt/Qwen2.5-7B-Instruct", "messages": [ {"role": "system", "content": "You are a friendly AI assistant."}, {"role": "user", "content": "Introduce deep learning."} ]}'For more information about using vLLM, see vLLM.
Known issues
The deepgpu-comfyui plugin accelerates Wanx model video generation but currently supports only the GN8IS, G49E, and G59 GPU types.