All Products
Search
Document Center

Container Compute Service:inference-nv-pytorch 25.12

Last Updated:May 10, 2026

This document contains the release notes for inference-nv-pytorch 25.12.

Main features and bug fixes

Main features

  • This release provides images for two CUDA versions:

    • The CUDA 12.8 image supports only the amd64 architecture.

    • The CUDA 13.0 image supports the amd64 and aarch64 architectures.

  • The PyTorch version is upgraded to 2.9.0 for vLLM images and to 2.9.1 for SGLang images.

  • For the CUDA 12.8 image, deepgpu-comfyui is upgraded to 1.3.2, and the deepgpu-torch optimization component is upgraded to 0.1.12+torch2.9.0cu128.

  • For both the CUDA 12.8 and 13.0 images, vLLM is upgraded to v0.12.0 and SGLang is upgraded to v0.5.6.post2.

Bug fixes

No bug fixes are included in this release.

Contents

Image name

inference-nv-pytorch

Tag

25.12-vllm0.12.0-pytorch2.9-cu128-20251215-serverless

25.12-sglang0.5.6.post2-pytorch2.9-cu128-20251215-serverless

25.12-vllm0.12.0-pytorch2.9-cu130-20251215-serverless

25.12-sglang0.5.6.post2-pytorch2.9-cu130-20251215-serverless

Supported architecture

amd64

amd64

amd64

aarch64

amd64

aarch64

Use case

large model inference

large model inference

large model inference

large model inference

large model inference

large model inference

Framework

pytorch

pytorch

pytorch

pytorch

pytorch

pytorch

Requirements

NVIDIA Driver release >= 570

NVIDIA Driver release >= 570

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

System components

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu128

  • CUDA 12.8

  • diffusers 0.36.0

  • deepgpu-comfyui 1.3.2

  • deepgpu-torch 0.1.12+torch2.9.0cu128

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • triton 3.5.0

  • torchaudio 2.9.0+cu128

  • torchvision 0.24.0+cu128

  • vllm 0.12.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu128

  • CUDA 12.8

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 2.0.0

  • deepgpu-comfyui 1.3.2

  • deepgpu-torch 0.1.12+torch2.9.0cu128

  • flash_attn 2.8.3

  • flash_mla 1.0.0+1408756

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.1

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • Triton 3.5.0

  • torchaudio 2.9.0+cu130

  • torchvision 0.24.0+cu130

  • vllm 0.12.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • transformers 4.57.1

  • ray 2.53.0

  • vllm 0.12.0

  • triton 3.5.0

  • torchaudio 2.9.0

  • torchvision 0.24.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 2.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord2 2.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • transformers 4.57.1

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • xfuser 0.4.5

Asset

Public images

CUDA 12.8 Asset

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-vllm0.12.0-pytorch2.9-cu128-20251215-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-sglang0.5.6.post2-pytorch2.9-cu128-20251215-serverless

CUDA 13.0 Asset

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-vllm0.12.0-pytorch2.9-cu130-20251215-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-sglang0.5.6.post2-pytorch2.9-cu130-20251215-serverless

VPC images

To quickly pull ACS AI container images within a Virtual Private Cloud (VPC), replace the image Asset URI egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/{image:tag} with acs-registry-vpc.{region-id}.cr.aliyuncs.com/egslingjun/{image:tag}.

  • {region-id}: The region ID for an ACS product in a supported region. For example: cn-beijing and cn-wulanchabu.

  • {image:tag}: The name and tag of the AI container image. For example: inference-nv-pytorch:25.10-vllm0.11.0-pytorch2.8-cu128-20251028-serverless and training-nv-pytorch:25.10-serverless.

Note

These images are for ACS and multi-tenant Lingjun services. Do not use these images for single-tenant Lingjun services.

Driver Requirements

  • CUDA 12.8: NVIDIA Driver release >= 570

  • CUDA 13.0: NVIDIA Driver release >= 580

Quick Start

The following example shows how to pull the inference-nv-pytorch image using Docker and test the inference service with the Qwen2.5-7B-Instruct model.

Note

To use the inference-nv-pytorch image in ACS, select the image from the Artifact Center page when you create a workload in the console. Alternatively, specify the image reference in a YAML file. For more information, see the following topics about building model inference services with ACS GPU resources:

  1. Pull the inference container image.

    docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  2. Download the open-source model from ModelScope.

    pip install modelscope
    cd /mnt
    modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-Instruct
  3. Run the following command to enter the container.

    docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864  \
    -v /mnt/:/mnt/ \
    egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  4. Test the vLLM conversational inference feature.

    1. Start the server-side service.

      python3 -m vllm.entrypoints.openai.api_server \
      --model /mnt/Qwen2.5-7B-Instruct \
      --trust-remote-code --disable-custom-all-reduce \
      --tensor-parallel-size 1
    2. Test on the client side.

      curl http://localhost:8000/v1/chat/completions \
          -H "Content-Type: application/json" \
          -d '{
          "model": "/mnt/Qwen2.5-7B-Instruct",  
          "messages": [
          {"role": "system", "content": "You are a friendly AI assistant."},
          {"role": "user", "content": "Tell me about deep learning."}
          ]}'

      For more information about vLLM, see vLLM.

Known issues

  • The deepgpu-comfyui plugin accelerates video generation for Wanx models. It currently supports only GN8IS, G49E, and G59.