Tous les produits
Search
Centre de documentation

Container Compute Service:inference-nv-pytorch 25.12

Dernière mise à jour :Aug 12, 2026

Ce document présente les notes de version d'inference-nv-pytorch 25.12.

Nouvelles fonctionnalités et corrections de bugs

Nouvelles fonctionnalités

  • Cette version propose des images pour deux versions de CUDA :

    • L'image CUDA 12.8 prend en charge uniquement l'architecture amd64.

    • L'image CUDA 13.0 prend en charge les architectures amd64 et aarch64 .

  • PyTorch passe en version 2.9.0 pour les images vLLM et en version 2.9.1 pour les images SGLang.

  • Pour l'image CUDA 12.8, deepgpu-comfyui passe en version 1.3.2 et le composant d'optimisation deepgpu-torch en version 0.1.12+torch2.9.0cu128.

  • Pour les images CUDA 12.8 et 13.0, vLLM passe en version v0.12.0 et SGLang en version v0.5.6.post2.

Corrections de bugs

Cette version n'inclut aucune correction de bug.

Contenu

Nom de l'image

inference-nv-pytorch

Tag

25.12-vllm0.12.0-pytorch2.9-cu128-20251215-serverless

25.12-sglang0.5.6.post2-pytorch2.9-cu128-20251215-serverless

25.12-vllm0.12.0-pytorch2.9-cu130-20251215-serverless

25.12-sglang0.5.6.post2-pytorch2.9-cu130-20251215-serverless

Architecture prise en charge

amd64

amd64

amd64

aarch64

amd64

aarch64

Cas d'utilisation

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Inférence de grands modèles

Framework

pytorch

pytorch

pytorch

pytorch

pytorch

pytorch

Prérequis

NVIDIA Driver release >= 570

NVIDIA Driver release >= 570

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

NVIDIA Driver release >= 580

Composants système

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu128

  • CUDA 12.8

  • diffusers 0.36.0

  • deepgpu-comfyui 1.3.2

  • deepgpu-torch 0.1.12+torch2.9.0cu128

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • triton 3.5.0

  • torchaudio 2.9.0+cu128

  • torchvision 0.24.0+cu128

  • vllm 0.12.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu128

  • CUDA 12.8

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 2.0.0

  • deepgpu-comfyui 1.3.2

  • deepgpu-torch 0.1.12+torch2.9.0cu128

  • flash_attn 2.8.3

  • flash_mla 1.0.0+1408756

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.1

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • xfuser 0.4.5

  • ljperf 0.1.0+477686c5

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • Triton 3.5.0

  • torchaudio 2.9.0+cu130

  • torchvision 0.24.0+cu130

  • vllm 0.12.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.0+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • transformers 4.57.1

  • ray 2.53.0

  • vllm 0.12.0

  • triton 3.5.0

  • torchaudio 2.9.0

  • torchvision 0.24.0

  • xfuser 0.4.5

  • xgrammar 0.1.27

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord 0.6.0

  • decord2 2.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • ray 2.52.1

  • transformers 4.57.3

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1+cu130

  • xfuser 0.4.5

  • ljperf 0.1.0+d0e4a408

  • Ubuntu 24.04

  • Python 3.12

  • Torch 2.9.1+cu130

  • CUDA 13.0.2

  • diffusers 0.36.0

  • decord2 2.0.0

  • flash_attn 2.8.3

  • flashinfer-python 0.5.3

  • imageio 2.37.2

  • imageio-ffmpeg 0.6.0

  • transformers 4.57.1

  • sgl-kernel 0.3.19

  • sglang 0.5.6.post2

  • xgrammar 0.1.27

  • triton 3.5.1

  • torchao 0.9.0

  • torchaudio 2.9.1

  • torchvision 0.24.1

  • xfuser 0.4.5

Ressources

Images publiques

Ressources CUDA 12.8

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-vllm0.12.0-pytorch2.9-cu128-20251215-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-sglang0.5.6.post2-pytorch2.9-cu128-20251215-serverless

Ressources CUDA 13.0

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-vllm0.12.0-pytorch2.9-cu130-20251215-serverless

  • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.12-sglang0.5.6.post2-pytorch2.9-cu130-20251215-serverless

Images VPC

Remarque

Ces images sont destinées aux services ACS et Lingjun multilocataires. Ne les utilisez pas pour les services Lingjun monolocataires.

Prérequis du pilote

  • CUDA 12.8 : NVIDIA Driver release >= 570

  • CUDA 13.0 : NVIDIA Driver release >= 580

Démarrage rapide

L'exemple suivant montre comment récupérer l'image inference-nv-pytorch via Docker et tester le service d'inférence avec le modèle Qwen2.5-7B-Instruct.

Remarque

Pour utiliser l'image inference-nv-pytorch dans ACS, sélectionnez-la depuis la page Artifact Center lors de la création d'une charge de travail dans la console. Vous pouvez également spécifier la référence de l'image dans un fichier YAML. Pour plus d'informations, consultez les rubriques suivantes sur la création de services d'inférence de modèles avec des ressources GPU ACS :

  1. Récupérez l'image du conteneur d'inférence.

    docker pull egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  2. Téléchargez le modèle open source depuis ModelScope.

    pip install modelscope
    cd /mnt
    modelscope download --model Qwen/Qwen2.5-7B-Instruct --local_dir ./Qwen2.5-7B-Instruct
  3. Exécutez la commande suivante pour accéder au conteneur.

    docker run -it --rm --gpus all --network=host --privileged --init --ipc=host \
    --ulimit memlock=-1 --ulimit stack=67108864  \
    -v /mnt/:/mnt/ \
    egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:[tag]
  4. Testez la fonctionnalité d'inférence conversationnelle vLLM.

    1. Démarrez le service côté serveur.

      python3 -m vllm.entrypoints.openai.api_server \
      --model /mnt/Qwen2.5-7B-Instruct \
      --trust-remote-code --disable-custom-all-reduce \
      --tensor-parallel-size 1
    2. Testez le service côté client.

      curl http://localhost:8000/v1/chat/completions \
          -H "Content-Type: application/json" \
          -d '{
          "model": "/mnt/Qwen2.5-7B-Instruct",  
          "messages": [
          {"role": "system", "content": "You are a friendly AI assistant."},
          {"role": "user", "content": "Tell me about deep learning."}
          ]}'

      Pour plus d'informations sur vLLM, consultez vLLM.

Problèmes connus

  • Le plugin deepgpu-comfyui accélère la génération vidéo pour les modèles Wanx. Il ne prend actuellement en charge que les instances GN8IS, G49E et G59.