This topic provides the release notes for inference-nv-pytorch 26.06.
Main features and bug fixes
Main features
In the vLLM image, vLLM is upgraded to v0.22.0.
In the SGLang image, SGLang is upgraded to v0.5.12.post1.
The vLLM and SGLang images now support amd64 and aarch64 architectures.
Bug fixes
None
Contents
Image name | inference-nv-pytorch | |||
Tag | 26.06-vllm0.22.0-pytorch2.11-cu130-20260611-serverless | 26.06-sglang0.5.12.post1-pytorch2.11-cu130-20260611-serverless | ||
Supported architecture | amd64 | aarch64 | amd64 | aarch64 |
Use case | large model inference | large model inference | large model inference | large model inference |
Framework | pytorch | pytorch | pytorch | pytorch |
Requirements | NVIDIA Driver release >= 580 | NVIDIA Driver release >= 580 | NVIDIA Driver release >= 580 | NVIDIA Driver release >= 580 |
System components |
|
|
|
|
Asset
Public image
CUDA 13.0 asset
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.06-vllm0.22.0-pytorch2.11-cu130-20260611-serverless
egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.06-sglang0.5.12.post1-pytorch2.11-cu130-20260622-serverless
VPC image
Replace the public network registry host in your image URI with the region-specific VPC endpoint.
URI component | Public network | IN-VPC |
Registry host |
|
|
Repository |
|
|
Image and tag |
|
|
Replace {region-id} with the ID of the region where your ACS service runs. For example:
Region | Region ID |
China (Beijing) |
|
China (Ulanqab) |
|
For the full list of supported regions, see Regions.
Example {image:tag} values:
inference-nv-pytorch:25.10-vllm0.11.0-pytorch2.8-cu128-20251028-serverlesstraining-nv-pytorch:25.10-serverless
These images are for ACS and EGS multi-tenant. Do not use them in EGS dedicated environments.
Driver requirements
CUDA 13.0: NVIDIA Driver release >= 580
Quick start
This example shows how to pull the inference-nv-pytorch image using Docker and test the inference service with the Qwen2.5-7B-Instruct model.
To use the inference-nv-pytorch image in ACS, select the image from the Artifacts Center on the Create Workload page in the console. You can also specify the image reference in a YAML file. For more information, see the following topics about building model inference services with ACS GPU resources:
Pull the inference container image.
docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]Download the open-source model from ModelScope.
pip install modelscope cd /mnt modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-InstructRun the following command to enter the container.
docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \ --ulimit memlock=-1 --ulimit stack=67108864 \ -v /mnt/:/mnt/ \ egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]Run an inference test for the vLLM conversational feature.
Start the server service.
python3 -m vllm.entrypoints.openai.api_server \ --model /mnt/Qwen2.5-7B-Instruct \ --trust-remote-code --disable-custom-all-reduce \ --tensor-parallel-size 1Test on the client.
curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "/mnt/Qwen2.5-7B-Instruct", "messages": [ {"role": "system", "content": "You are a friendly AI assistant."}, {"role": "user", "content": "Introduce deep learning."} ]}'For more information about how to use vLLM, see vLLM.
Known issues
The images in this release do not support the deepgpu-comfyui plugin.