This document provides the release notes for inference-nv-pytorch 26.03.
Main features and bug fixes
Main features
This release provides images for CUDA 12.8 and CUDA 13.0:
The CUDA 12.8 image supports only the amd64 architecture.
The CUDA 13.0 image supports the amd64 and aarch64 architectures.
This release upgrades vLLM to v0.17.1, adding support for the Qwen3.5 model.
Bug fixes
None.
Contents
Image name | inference-nv-pytorch | ||
Tag | 26.03-vllm0.17.1-pytorch2.10-cu128-20260317-serverless | 26.03-vllm0.17.1-pytorch2.10-cu130-20260317-serverless | |
Supported architecture | amd64 | amd64 | aarch64 |
Use case | large model inference | large model inference | large model inference |
Framework | PyTorch | PyTorch | PyTorch |
Requirements | NVIDIA Driver release >= 570 | NVIDIA Driver release >= 580 | NVIDIA Driver release >= 580 |
System components |
|
|
|
Assets
Public image
CUDA 12.8 asset
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.03-vllm0.17.1-pytorch2.10-cu128-20260317-serverless
CUDA 13.0 asset
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.03-vllm0.17.1-pytorch2.10-cu130-20260317-serverless
VPC image
All existing and new ACS AI container images in the egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun repository support IN-VPC pulling.
Replace the public network registry host in your image URI with the region-specific VPC endpoint.
URI component | Public network | IN-VPC |
Registry host |
|
|
Repository |
|
|
Image and tag |
|
|
Replace {region-id} with the ID of the region where your ACS service runs. For example:
Region | Region ID |
China (Beijing) |
|
China (Ulanqab) |
|
For the full list of supported regions, see Regions.
Example {image:tag} values:
inference-nv-pytorch:25.10-vllm0.11.0-pytorch2.8-cu128-20251028-serverlesstraining-nv-pytorch:25.10-serverless
These images are for AI Computing Service (ACS) and Lingjun multi-tenant environments, and are not supported in Lingjun single-tenant environments.
Driver requirements
CUDA 12.8: NVIDIA Driver release >= 570
CUDA 13.0: NVIDIA Driver release >= 580
Quick start
The following example shows how to pull the inference-nv-pytorch image using Docker and test the inference service with the Qwen2.5-7B-Instruct model.
To use the inference-nv-pytorch image in AI Computing Service (ACS), select it from the artifact repository page when you create a workload in the console, or specify the image reference in a YAML file. For more information, see the following topics on building model inference services with ACS GPU computing power:
Pull the inference container image.
docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]Download the open-source model from ModelScope.
pip install modelscope cd /mnt modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-InstructRun the following command to enter the container.
docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \ --ulimit memlock=-1 --ulimit stack=67108864 \ -v /mnt/:/mnt/ \ egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]Run an inference test to verify the vLLM chat completion feature.
Start the server-side service.
python3 -m vllm.entrypoints.openai.api_server \ --model /mnt/Qwen2.5-7B-Instruct \ --trust-remote-code --disable-custom-all-reduce \ --tensor-parallel-size 1Run a test on the client side.
curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "/mnt/Qwen2.5-7B-Instruct", "messages": [ {"role": "system", "content": "You are a friendly AI assistant."}, {"role": "user", "content": "Introduce deep learning."} ]}'For more information, see the vLLM documentation.
Known issues
This image does not support the deepgpu-comfyui plugin.