All Products
Search
Document Center

Container Compute Service:Build a DeepSeek distilled model inference service using ACS GPU computing power

Last Updated:Aug 28, 2026

Container Compute Service (ACS) lets you run GPU-accelerated LLM inference without managing the underlying hardware or GPU nodes. All configurations work out of the box, and billing is pay-as-you-go. ACS is ideal for large language model (LLM) inference tasks and helps reduce inference costs. This topic walks you through deploying a production-ready DeepSeek distilled model inference service on ACS using the vLLM framework.

Background

DeepSeek-R1

DeepSeek-R1 is DeepSeek's first-generation reasoning model, trained with large-scale reinforcement learning to improve LLM reasoning performance. It outperforms other closed-source models in mathematical reasoning and programming competitions, and its performance approaches or surpasses the OpenAI o1 series on certain tasks. DeepSeek-R1 also delivers strong results across knowledge-intensive tasks such as writing and Q&A.

DeepSeek also distills its reasoning capabilities into smaller models based on the Qwen and Llama architectures. The 14B distilled model outperforms the open-source QwQ-32B model, while the 32B and 70B distilled models set new benchmarks. For more information, see the DeepSeek AI GitHub repository.

vLLM

vLLM is a high-performance LLM inference framework that supports most widely used LLMs, including the Qwen series. It uses PagedAttention, continuous batching, and model quantization to maximize inference throughput. For more information, see the vLLM GitHub repository.

Arena

Arena is a lightweight client for managing Kubernetes-based machine learning tasks. It covers the full ML lifecycle—data preparation, model development, training, and prediction—and improves data scientists' productivity. Arena deeply integrates with Alibaba Cloud infrastructure services, including GPU sharing and Cloud Parallel File Storage (CPFS), and can run Alibaba Cloud-optimized deep learning frameworks to maximize the performance and cost benefits of Alibaba Cloud heterogeneous devices. For more information, see the Arena GitHub repository.

Prerequisites

Before you begin, make sure you have:

  • Assigned the system default role to the service account (required when you use Alibaba Cloud for the first time) so ACS can call dependent services such as Elastic Compute Service (ECS), Object Storage Service (OSS), Apsara File Storage NAS, Cloud Parallel File Storage (CPFS), and Server Load Balancer (SLB), create clusters, and save logs. ACS can use these capabilities only after this role is correctly granted. For details, see Get started with Container Compute Service.

  • An ACS cluster in a region and zone that provides GPU resources. See Create an ACS cluster.

  • You have connected to a Kubernetes cluster using kubectl.

  • The Arena client installed. See Configure the Arena client.

Choose a GPU instance specification

GPU memory during inference is consumed primarily by model weights, KV cache, and runtime buffers. Use the following formula to estimate the minimum GPU memory required:

GPU memory = number of model parameters × bytes per parameter

For a 7B FP16 model: 7 × 10⁹ × 2 bytes ≈ 13.04 GiB. Because you must also consider the KV cache size required during computation, GPU utilization, and the buffer reserved for them, the suggested specification for the 7B model is 1 GPU with 24 GiB of memory, 8 vCPUs, and 32 GiB of RAM.

The following table lists suggested specifications for each supported DeepSeek distilled model. For the full list of GPU models and specifications available in ACS, see GPU models and specifications. For billing details, see Billing overview.

Model

Size

Model size (GB)

Specification recommended

vCPUs

RAM

GPU memory

DeepSeek-R1-Distill-Qwen-1.5B

1.5B

3.55 GB

4 or 6

30 GiB

24 GiB

DeepSeek-R1-Distill-Qwen-7B

7B

15.23 GB

6 or 8

32 GiB

24 GiB

DeepSeek-R1-Distill-Llama-8B

8B

16.06 GB

6 or 8

32 GiB

24 GiB

DeepSeek-R1-Distill-Qwen-14B

14B

29.54 GB

8+

64 GiB

48 GiB

DeepSeek-R1-Distill-Qwen-32B

32B

74.32 GB

8+

128 GiB

96 GiB

DeepSeek-R1-Distill-Llama-70B

70B

140.56 GB

12+

128 GiB

192 GiB

Note

Procedures

Step 1: Prepare the model files

Note

Downloading and uploading the model typically takes 1 to 2 hours. To save time, submit a ticket to have the model files copied directly to your OSS bucket.

  1. Download the DeepSeek-R1-Distill-Qwen-7B model from ModelScope using git-lfs.

    Note

    Check whether git-lfs is installed. If not, install it by running yum install git-lfs or apt-get install git-lfs. See Install Git Large File Storage.

    git lfs install
    GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B.git
    cd DeepSeek-R1-Distill-Qwen-7B/
    git lfs pull
  2. Create an OSS directory and upload the model files.

    Note

    To install ossutil, see Install ossutil.

    ossutil mkdir oss://<your-bucket-name>/models/DeepSeek-R1-Distill-Qwen-7B
    ossutil cp -r ./DeepSeek-R1-Distill-Qwen-7B oss://<your-bucket-name>/models/DeepSeek-R1-Distill-Qwen-7B
  3. Create a PV and a PVC. Configure a persistent volume (PV) named llm-modeland a persistent volume claim (PVC) for the target cluster. For a full walkthrough, see Use a static OSS volume.

    Use the following parameters when creating the PV:

    Parameter

    Value

    PV type

    OSS

    Volume name

    llm-model

    Access certificate

    The AccessKey ID and AccessKey secret for your OSS bucket

    Bucket ID

    The OSS bucket you created in the previous step

    OSS path

    /models/DeepSeek-R1-Distill-Qwen-7B

    Use the following parameters when creating the PVC:

    Parameter

    Value

    PVC type

    OSS

    Name

    llm-model

    Allocation mode

    Existing volumes

    Existing volumes

    Select the PV created above

    Yaml example:

    apiVersion: v1
    kind: Secret
    metadata:
      name: oss-secret
    stringData:
      akId: <your-oss-ak>       # AccessKey ID for the OSS bucket
      akSecret: <your-oss-sk>   # AccessKey secret for the OSS bucket
    ---
    apiVersion: v1
    kind: PersistentVolume
    metadata:
      name: llm-model
      labels:
        alicloud-pvname: llm-model
    spec:
      capacity:
        storage: 30Gi
      accessModes:
        - ReadOnlyMany
      persistentVolumeReclaimPolicy: Retain
      csi:
        driver: ossplugin.csi.alibabacloud.com
        volumeHandle: llm-model
        nodePublishSecretRef:
          name: oss-secret
          namespace: default
        volumeAttributes:
          bucket: <your-bucket-name>       # Name of the OSS bucket
          url: <your-bucket-endpoint>      # Endpoint, for example, oss-cn-hangzhou-internal.aliyuncs.com
          otherOpts: "-o umask=022 -o max_stat_cache_size=0 -o allow_other"
          path: <your-model-path>          # Model path, for example, /models/DeepSeek-R1-Distill-Qwen-7B/
    ---
    apiVersion: v1
    kind: PersistentVolumeClaim
    metadata:
      name: llm-model
    spec:
      accessModes:
        - ReadOnlyMany
      resources:
        requests:
          storage: 30Gi
      selector:
        matchLabels:
          alicloud-pvname: llm-model

Step 2: Deploy the model

  1. Run the following command to deploy the vLLM inference service. The service exposes an OpenAI-compatible HTTP API

    The --data flag mounts the model PVC to /model/DeepSeek-R1-Distill-Qwen-7B inside the container. The --max-model-len flag sets the maximum context length in tokens—increasing it can improve the model's conversation quality but may consume more GPU memory.

    Note
    • Replace <example-model> in --label=alibabacloud.com/gpu-model-series=<example-model> with the actual GPU model series supported by your ACS cluster. Submit a ticket to get the list of available GPU model series.

    • egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/{image:tag} is a public registry address. To reduce image pull time, we recommend using VPC to accelerate AI container image pulls.

    arena serve custom \
    --name=deepseek-r1 \
    --version=v1 \
    --gpus=1 \
    --cpu=8 \
    --memory=32Gi \
    --replicas=1 \
    --label=alibabacloud.com/compute-class=gpu \
    --label=alibabacloud.com/gpu-model-series=<example-model> \
    --restful-port=8000 \
    --readiness-probe-action="tcpSocket" \
    --readiness-probe-action-option="port: 8000" \
    --readiness-probe-option="initialDelaySeconds: 30" \
    --readiness-probe-option="periodSeconds: 30" \
    --image=egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.02-vllm0.7.2-sglang0.4.3.post2-pytorch2.5-cuda12.4-20250305-serverless \
    --data=llm-model:/models/DeepSeek-R1-Distill-Qwen-7B \
    "vllm serve /models/DeepSeek-R1-Distill-Qwen-7B --port 8000 --trust-remote-code --served-model-name deepseek-r1 --max-model-len 32768 --gpu-memory-utilization 0.95 --enforce-eager"

    The expected output is:

    service/deepseek-r1-v1 created
    deployment.apps/deepseek-r1-v1-custom-serving created
    INFO[0004] The Job deepseek-r1 has been submitted successfully
    INFO[0004] You can run `arena serve get deepseek-r1 --type custom-serving -n default` to check the job status

    The following table describes the key parameters.

    Parameter

    Description

    --name

    Name of the inference service

    --version

    Version of the inference service

    --gpus

    Number of GPUs per replica

    --cpu

    Number of vCPUs per replica

    --memory

    Amount of memory per replica

    --replicas

    Number of replicas

    --label

    Labels that specify ACS GPU compute power. Set alibabacloud.com/compute-class=gpu and alibabacloud.com/gpu-model-series=<example-model>

    --restful-port

    Port exposed by the inference service

    --readiness-probe-action

    Connection type for readiness probes. Valid values: httpGet, exec, grpc, tcpSocket

    --readiness-probe-action-option

    Connection method for readiness probes

    --readiness-probe-option

    Readiness probe configuration

    --image

    Container image for the inference service

    --data

    Mounts a PVC into the container. Format: <pvc-name>:<mount-path>. Run arena data list to list available PVCs

  2. To check the deployment status, run:

    arena serve get deepseek-r1

    Expected output:

    Name:       deepseek-r1
    Namespace:  default
    Type:       Custom
    Version:    v1
    Desired:    1
    Available:  1
    Age:        6h
    Address:    10.0.78.27
    Port:       RESTFUL:8000
    GPU:        1
    
    Instances:
      NAME                                            STATUS   AGE  READY  RESTARTS  GPU  NODE
      ----                                            ------   ---  -----  --------  ---  ----
      deepseek-r1-v1-custom-serving-54d579d994-dqwxz  Running  1h   1/1    0         1    virtual-kubelet-cn-hangzhou-b

Step 3: Verify the inference service

  1. Use kubectl port-forward to set up port forwarding from local machine to the inference service.

    Note

    The port forwarding established by kubectl port-forward does not provide production-grade reliability, security, or scalability. Therefore, it is intended only for development and debugging purposes and is not suitable for production environments. For more information about production-ready networking solutions in Kubernetes clusters, see Ingress management.

    kubectl port-forward svc/deepseek-r1-v1 8000:8000

    Expected output:

    Forwarding from 127.0.0.1:8000 -> 8000
    Forwarding from [::1]:8000 -> 8000
  2. Send a test request to the OpenAI-compatible chat completions endpoint.

    curl http://localhost:8000/v1/chat/completions \
      -H "Content-Type: application/json" \
      -d '{
        "model": "deepseek-r1",
        "messages": [
          {
            "role": "user",
            "content": "Write a letter to my daughter from the future 2035 and tell her to study science and technology well, be the master of science and technology, and promote the development of science and technology and economy. She is now in grade 3."
          }
        ],
        "max_tokens": 1024,
        "temperature": 0.7,
        "top_p": 0.9,
        "seed": 10
      }'

    Expected output:

    {
      "id": "chatcmpl-53613fd815da46df92cc9b92cd156146",
      "object": "chat.completion",
      "created": 1739261570,
      "model": "deepseek-r1",
      "choices": [
        {
          "index": 0,
          "message": {
            "role": "assistant",
            "content": "<think>...</think>\n\nDear Future 2035: ..."
          },
          "finish_reason": "stop"
        }
      ],
      "usage": {
        "prompt_tokens": 40,
        "completion_tokens": 994,
        "total_tokens": 1034
      }
    }

(Optional) Step 4: Clean up

Delete the inference service and storage resources when they are no longer needed.

  1. Delete the inference service.

    arena serve delete deepseek-r1

    Expected output:

    INFO[0007] The serving job deepseek-r1 with version v1 has been deleted successfully
  2. Delete the PVC and PV.

    kubectl delete pvc llm-model
    kubectl delete pv llm-model

    Expected output:

    persistentvolumeclaim "llm-model" deleted
    persistentvolume "llm-model" deleted

What's next