Tous les produits
Search
Centre de documentation

Container Compute Service:inference-nv-pytorch 26.02

Dernière mise à jour :Aug 12, 2026

Ce document présente les notes de version d'inference-nv-pytorch 26.02.

Nouveautés et correctifs

Nouvelles fonctionnalités

  • Cette version propose des images pour deux versions de CUDA : CUDA 12.8 et CUDA 13.0.

    • L'image CUDA 12.8 prend uniquement en charge l'architecture amd64.

    • L'image CUDA 13.0 est compatible avec les architectures amd64 et aarch64.

  • Dans l'image vLLM, Torch passe à la version 2.10.0 et vLLM à la version v0.15.1.

  • Dans l'image SGLang, Torch passe à la version 2.10.0 et SGLang à la version v0.5.9.

Correctifs de bugs

Aucun correctif de bug n'est inclus dans cette version.

Contenu

Nom de l'image

inference-nv-pytorch

Tag

26.02-vllm0.15.1-pytorch2.10-cu128-20260211-serverless

26.02-sglang0.5.9-pytorch2.10-cu128-20260227-serverless

26.02-vllm0.15.1-pytorch2.10-cu130-20260211-serverless

26.02-sglang0.5.9-pytorch2.10-cu130-20260227-serverless

Architecture prise en charge

amd64

amd64

amd64

aarch64

amd64

aarch64

Cas d'utilisation

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Framework

pytorch

pytorch

pytorch

pytorch

pytorch

pytorch

Prérequis

NVIDIA Driver release >= 570

NVIDIA Driver release >= 570

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

Composants système

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0

  • CUDA 12.8

  • NCCL 2.28.9

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flash_attn_3 3.0.0

  • flashinfer-python 0.6.1

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.0

  • transformers 4.57.6

  • triton 3.6.0

  • torchaudio 2.10.0

  • torchvision 0.25.0

  • vllm 0.15.1

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0

  • CUDA 12.8

  • NCCL 2.28.9

  • torchaudio 2.10.0

  • torchvision 0.25.0

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flash_attn_3 3.0.0

  • flashinfer-python 0.6.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.54.0

  • transformers 4.57.1

  • sgl-kernel 0.3.21

  • sglang 0.5.9

  • xgrammar 0.1.27

  • triton 3.6.0

  • torchao 0.9.0

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flash_attn_3 3.0.0

  • flashinfer-python 0.6.1

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.53.0

  • transformers 4.57.6

  • triton 3.6.0

  • torchaudio 2.10.0+cu130

  • torchvision 0.25.0+cu130

  • vllm 0.15.1

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.6.1

  • transformers 4.57.6

  • ray 2.53.0

  • vllm 0.15.1

  • triton 3.6.0

  • torchaudio 2.10.0+cu130

  • torchvision 0.25.0+cu130

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.6.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.54.0

  • transformers 4.57.1

  • sgl-kernel 0.3.21

  • sglang 0.5.9

  • xgrammar 0.1.27

  • triton 3.6.0

  • torchao 0.9.0

  • torchaudio 2.10.0+cu130

  • torchvision 0.25.0+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.10.0+cu130

  • CUDA 13.0.2

  • NCCL 2.28.9

  • diffusers 0.36.0

  • decord2 3.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.6.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.54.0

  • transformers 4.57.1

  • sgl-kernel 0.3.21

  • sglang 0.5.9

  • xgrammar 0.1.27

  • triton 3.6.0

  • torchao 0.9.0

  • torchaudio 2.10.0+cu130

  • torchvision 0.25.0+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

Ressources

Image publique

Ressources CUDA 12.8

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-vllm0.15.1-pytorch2.10-cu128-20260211-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-sglang0.5.9-pytorch2.10-cu128-20260227-serverless

Ressources CUDA 13.0

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-vllm0.15.1-pytorch2.10-cu130-20260211-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:26.02-sglang0.5.9-pytorch2.10-cu130-20260227-serverless

Image VPC

Remarque

Cette image est destinée aux produits multi-locataires ACS et Lingjun. Elle n'est pas prise en charge dans les scénarios Lingjun mono-locataire.

Prérequis relatifs aux pilotes

  • CUDA 12.8 : NVIDIA Driver release >= 570

  • CUDA 13.0 : NVIDIA Driver release >= 580

Démarrage rapide

L'exemple suivant montre comment récupérer l'image inference-nv-pytorch avec Docker et tester le service d'inférence avec le modèle Qwen2.5-7B-Instruct.

Remarque

Pour utiliser l'image inference-nv-pytorch dans ACS, sélectionnez-la sur la page Artifacts Center lors de la création d'une charge de travail dans la console, ou spécifiez sa référence dans un fichier YAML. Pour plus d'informations, consultez les rubriques suivantes détaillant la création d'un service d'inférence de modèle avec le calcul GPU ACS :

  1. Récupérez l'image du conteneur d'inférence.

    docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  2. Téléchargez le modèle open source depuis ModelScope.

    pip install modelscope
    cd /mnt
    modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-Instruct
  3. Exécutez la commande suivante pour démarrer le conteneur et y accéder.

    docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864  \
    -v /mnt/:/mnt/ \
    egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  4. Testez la fonctionnalité conversationnelle de vLLM.

    1. Démarrez le service côté serveur.

      python3 -m vllm.entrypoints.openai.api_server \
      --model /mnt/Qwen2.5-7B-Instruct \
      --trust-remote-code --disable-custom-all-reduce \
      --tensor-parallel-size 1
    2. Testez le service côté client.

      curl http://localhost:8000/v1/chat/completions \
          -H "Content-Type: application/json" \
          -d '{
          "model": "/mnt/Qwen2.5-7B-Instruct",  
          "messages": [
          {"role": "system", "content": "You are a helpful AI assistant."},
          {"role": "user", "content": "Introduce deep learning."}
          ]}'

      Pour plus d'informations sur l'utilisation de vLLM, consultez la documentation vLLM.

Problèmes connus

  • Les images de cette version ne prennent pas en charge le plug-in deepgpu-comfyui.