All Products
Search
Document Center

Container Compute Service:inference-nv-pytorch 26.01

Last Updated:Aug 28, 2026

The inference-nv-pytorch 26.01 release provides container images for large model inference on the ACS and Lingjun multi-tenant product formats. Use these release notes to choose an image Tag, confirm the system components that each image bundles, and check the NVIDIA driver release that each image requires.

Main features and bug fixes

Main features

  • Provides images for two CUDA versions: CUDA 12.8 and CUDA 13.0.

    • The CUDA 12.8 image supports only the amd64 architecture.

    • The CUDA 13.0 image supports both amd64 and aarch64 architectures.

  • In the CUDA 12.8 image, deepgpu-comfyui is upgraded to 1.4.1, and the deepgpu-torch optimization component is upgraded to 0.1.18+torch2.9.0cu128.

  • In both CUDA 12.8 and CUDA 13.0 images, the vLLM version is upgraded to v0.14.0, and the SGLang version is upgraded to v0.5.7.

Bug fixes

None.

Contents

Image name

inference-nv-pytorch

inference-nv-pytorch

inference-nv-pytorch

inference-nv-pytorch

inference-nv-pytorch

inference-nv-pytorch

Tag

26.01-vllm0.14.0-pytorch2.9-cu128-20260121-serverless

26.01-sglang0.5.7-pytorch2.9-cu128-20260113-serverless

26.01-vllm0.14.0-pytorch2.9-cu130-20260123-serverless

26.01-vllm0.14.0-pytorch2.9-cu130-20260123-serverless

26.01-sglang0.5.7-pytorch2.9-cu130-20260113-serverless

26.01-sglang0.5.7-pytorch2.9-cu130-20260113-serverless

Supported architecture

amd64

amd64

amd64

aarch64

amd64

aarch64

Application scenario

large model inference

large model inference

large model inference

large model inference

large model inference

large model inference

Framework

pytorch

pytorch

pytorch

pytorch

pytorch

pytorch

Requirements

NVIDIA Driver release >= 570

NVIDIA Driver release >= 570

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

System components

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1

  • CUDA 12.8

  • diffusers 0.36.0

  • deepgpu-comfyui 1.4.1

  • deepgpu-torch 0.1.18+torch2.9.0cu128

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.0

  • transformers 4.57.6

  • triton 3.5.1

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • vllm 0.14.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu128

  • CUDA 12.8

  • torchaudio 2.9.1+128

  • torchvision 0.24.1+128

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 3.0.0

  • deepgpu-comfyui 1.4.1

  • deepgpu-torch 0.1.18+torch2.9.0cu128

  • flash_attn 2.8.3

  • flash_mla 1.0.0+1408756

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.0

  • transformers 4.57.1

  • sgl-kernel 0.3.20

  • sglang 0.5.7

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.1

  • transformers 4.57.6

  • triton 3.5.0

  • torchaudio 2.9.1+cu130

  • torchvision 0.24.1+cu130

  • vllm 0.14.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • transformers 4.57.6

  • ray 2.53.0

  • vllm 0.14.0

  • triton 3.5.1

  • torchaudio 2.9.1+cu130

  • torchvision 2.9.1+cu130

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.0

  • transformers 4.57.1

  • sgl-kernel 0.3.20

  • sglang 0.5.7

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • transformers 4.57.1

  • sgl-kernel 0.3.20

  • sglang 0.5.7

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • xfuser 0.4.5

Driver requirements

  • CUDA 12.8: NVIDIA Driver release >= 570

  • CUDA 13.0: NVIDIA Driver release >= 580

Assets

Public images

CUDA 12.8 assets

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-vllm0.14.0-pytorch2.9-cu128-20260121-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-sglang0.5.7-pytorch2.9-cu128-20260113-serverless

CUDA 13.0 assets

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-vllm0.14.0-pytorch2.9-cu130-20260123-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.01-sglang0.5.7-pytorch2.9-cu130-20260113-serverless

VPC images

Replace the AI container image asset URI egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/{image:tag} specified in the YAML file of the ACS console with acs-registry-vpc.{region-id}.cr.aliyuncs.com/egslingjun/{image:tag} to quickly pull PG1 AI container images over the VPC.

  • Where {region-id} is the available region where your ACS is activated, such as cn-beijing and cn-wulanchabu.

  • {image:tag} is the name and tag of the image.

These images apply to the ACS product format and the Lingjun multi-tenant product format. They are not compatible with the Lingjun single-tenant product format. Do not use these images in a Lingjun single-tenant environment.

Quick start

The following example shows how to pull the inference-nv-pytorch image only by using Docker and test the inference service with the Qwen2.5-7B-Instruct model.

To use the inference-nv-pytorch image in ACS, select it from the Artifact Center page when you create a workload in the console, or specify the image reference in a YAML file. For more information, see the following topics about building model inference services and accelerating model workloads by using ACS GPU compute power:

  1. Pull the inference container image. Replace [tag] with an image Tag listed in the Contents table or the Public images section of this topic.

    docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  2. Download the open-source model from ModelScope.

    pip install modelscope
    cd /mnt
    modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-Instruct
  3. Run the following command to enter the container.

    docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864  \
    -v /mnt/:/mnt/ \
    egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  4. Run an inference test on the vLLM chat feature.

    1. Start the server-side service.

      python3 -m vllm.entrypoints.openai.api_server \
      --model /mnt/Qwen2.5-7B-Instruct \
      --trust-remote-code --disable-custom-all-reduce \
      --tensor-parallel-size 1
    2. Run the test on the client.

      curl http://localhost:8000/v1/chat/completions \
          -H "Content-Type: application/json" \
          -d '{
          "model": "/mnt/Qwen2.5-7B-Instruct",  
          "messages": [
          {"role": "system", "content": "You are a friendly AI assistant."},
          {"role": "user", "content": "Introduce deep learning."}
          ]}'

For more information about how to use vLLM, see the project repository.

Known issues

  • The deepgpu-comfyui plugin for accelerating Wanx model video generation currently supports only the GN8IS, G49E, and G59 instance types.