Container Compute Service (ACS) lets you run GPU-accelerated LLM inference without managing the underlying hardware or GPU nodes. All configurations work out of the box, and billing is pay-as-you-go. ACS is ideal for large language model (LLM) inference tasks and helps reduce inference costs. This topic walks you through deploying a production-ready DeepSeek distilled model inference service on ACS using the vLLM framework.
Background
DeepSeek-R1
vLLM
Arena
Prerequisites
Before you begin, make sure you have:
Assigned the system default role to the service account (required when you use Alibaba Cloud for the first time) so ACS can call dependent services such as Elastic Compute Service (ECS), Object Storage Service (OSS), Apsara File Storage NAS, Cloud Parallel File Storage (CPFS), and Server Load Balancer (SLB), create clusters, and save logs. ACS can use these capabilities only after this role is correctly granted. For details, see Get started with Container Compute Service.
An ACS cluster in a region and zone that provides GPU resources. See Create an ACS cluster.
The Arena client installed. See Configure the Arena client.
Choose a GPU instance specification
GPU memory during inference is consumed primarily by model weights, KV cache, and runtime buffers. Use the following formula to estimate the minimum GPU memory required:
GPU memory = number of model parameters × bytes per parameterFor a 7B FP16 model: 7 × 10⁹ × 2 bytes ≈ 13.04 GiB. Because you must also consider the KV cache size required during computation, GPU utilization, and the buffer reserved for them, the suggested specification for the 7B model is 1 GPU with 24 GiB of memory, 8 vCPUs, and 32 GiB of RAM.
The following table lists suggested specifications for each supported DeepSeek distilled model. For the full list of GPU models and specifications available in ACS, see GPU models and specifications. For billing details, see Billing overview.
Model | Size | Model size (GB) | Specification recommended | ||
vCPUs | RAM | GPU memory | |||
DeepSeek-R1-Distill-Qwen-1.5B | 1.5B | 3.55 GB | 4 or 6 | 30 GiB | 24 GiB |
DeepSeek-R1-Distill-Qwen-7B | 7B | 15.23 GB | 6 or 8 | 32 GiB | 24 GiB |
DeepSeek-R1-Distill-Llama-8B | 8B | 16.06 GB | 6 or 8 | 32 GiB | 24 GiB |
DeepSeek-R1-Distill-Qwen-14B | 14B | 29.54 GB | 8+ | 64 GiB | 48 GiB |
DeepSeek-R1-Distill-Qwen-32B | 32B | 74.32 GB | 8+ | 128 GiB | 96 GiB |
DeepSeek-R1-Distill-Llama-70B | 70B | 140.56 GB | 12+ | 128 GiB | 192 GiB |
Make sure the chosen specification complies with ACS pod specification adjustment logic.
By default, each ACS pod provides 30 GiB of free EphemeralStorage. The inference image inference-nv-pytorch:25.02-vllm0.7.2-sglang0.4.3.post2-pytorch2.5-cuda12.4-20250305-serverless used in this example is about 9.8 GiB. If you need more space, see Increase the temporary storage space size.
Procedures
Step 1: Prepare the model files
Downloading and uploading the model typically takes 1 to 2 hours. To save time, submit a ticket to have the model files copied directly to your OSS bucket.
Download the DeepSeek-R1-Distill-Qwen-7B model from ModelScope using git-lfs.
NoteCheck whether git-lfs is installed. If not, install it by running
yum install git-lfsorapt-get install git-lfs. See Install Git Large File Storage.git lfs install GIT_LFS_SKIP_SMUDGE=1 git clone https://www.modelscope.cn/deepseek-ai/DeepSeek-R1-Distill-Qwen-7B.git cd DeepSeek-R1-Distill-Qwen-7B/ git lfs pullCreate an OSS directory and upload the model files.
NoteTo install ossutil, see Install ossutil.
ossutil mkdir oss://<your-bucket-name>/models/DeepSeek-R1-Distill-Qwen-7B ossutil cp -r ./DeepSeek-R1-Distill-Qwen-7B oss://<your-bucket-name>/models/DeepSeek-R1-Distill-Qwen-7BCreate a PV and a PVC. Configure a persistent volume (PV) named
llm-modeland a persistent volume claim (PVC) for the target cluster. For a full walkthrough, see Use a static OSS volume.Use the following parameters when creating the PV:
Parameter
Value
PV type
OSS
Volume name
llm-modelAccess certificate
The AccessKey ID and AccessKey secret for your OSS bucket
Bucket ID
The OSS bucket you created in the previous step
OSS path
/models/DeepSeek-R1-Distill-Qwen-7BUse the following parameters when creating the PVC:
Parameter
Value
PVC type
OSS
Name
llm-modelAllocation mode
Existing volumes
Existing volumes
Select the PV created above
Yaml example:
apiVersion: v1 kind: Secret metadata: name: oss-secret stringData: akId: <your-oss-ak> # AccessKey ID for the OSS bucket akSecret: <your-oss-sk> # AccessKey secret for the OSS bucket --- apiVersion: v1 kind: PersistentVolume metadata: name: llm-model labels: alicloud-pvname: llm-model spec: capacity: storage: 30Gi accessModes: - ReadOnlyMany persistentVolumeReclaimPolicy: Retain csi: driver: ossplugin.csi.alibabacloud.com volumeHandle: llm-model nodePublishSecretRef: name: oss-secret namespace: default volumeAttributes: bucket: <your-bucket-name> # Name of the OSS bucket url: <your-bucket-endpoint> # Endpoint, for example, oss-cn-hangzhou-internal.aliyuncs.com otherOpts: "-o umask=022 -o max_stat_cache_size=0 -o allow_other" path: <your-model-path> # Model path, for example, /models/DeepSeek-R1-Distill-Qwen-7B/ --- apiVersion: v1 kind: PersistentVolumeClaim metadata: name: llm-model spec: accessModes: - ReadOnlyMany resources: requests: storage: 30Gi selector: matchLabels: alicloud-pvname: llm-model
Step 2: Deploy the model
Run the following command to deploy the vLLM inference service. The service exposes an OpenAI-compatible HTTP API
The
--dataflag mounts the model PVC to/model/DeepSeek-R1-Distill-Qwen-7Binside the container. The--max-model-lenflag sets the maximum context length in tokens—increasing it can improve the model's conversation quality but may consume more GPU memory.NoteReplace
<example-model>in--label=alibabacloud.com/gpu-model-series=<example-model>with the actual GPU model series supported by your ACS cluster. Submit a ticket to get the list of available GPU model series.egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/{image:tag}is a public registry address. To reduce image pull time, we recommend using VPC to accelerate AI container image pulls.
arena serve custom \ --name=deepseek-r1 \ --version=v1 \ --gpus=1 \ --cpu=8 \ --memory=32Gi \ --replicas=1 \ --label=alibabacloud.com/compute-class=gpu \ --label=alibabacloud.com/gpu-model-series=<example-model> \ --restful-port=8000 \ --readiness-probe-action="tcpSocket" \ --readiness-probe-action-option="port: 8000" \ --readiness-probe-option="initialDelaySeconds: 30" \ --readiness-probe-option="periodSeconds: 30" \ --image=egslingjun-registry.cn-wulanchabu.cr.aliyuncs.com/egslingjun/inference-nv-pytorch:25.02-vllm0.7.2-sglang0.4.3.post2-pytorch2.5-cuda12.4-20250305-serverless \ --data=llm-model:/models/DeepSeek-R1-Distill-Qwen-7B \ "vllm serve /models/DeepSeek-R1-Distill-Qwen-7B --port 8000 --trust-remote-code --served-model-name deepseek-r1 --max-model-len 32768 --gpu-memory-utilization 0.95 --enforce-eager"The expected output is:
service/deepseek-r1-v1 created deployment.apps/deepseek-r1-v1-custom-serving created INFO[0004] The Job deepseek-r1 has been submitted successfully INFO[0004] You can run `arena serve get deepseek-r1 --type custom-serving -n default` to check the job statusThe following table describes the key parameters.
Parameter
Description
--nameName of the inference service
--versionVersion of the inference service
--gpusNumber of GPUs per replica
--cpuNumber of vCPUs per replica
--memoryAmount of memory per replica
--replicasNumber of replicas
--labelLabels that specify ACS GPU compute power. Set
alibabacloud.com/compute-class=gpuandalibabacloud.com/gpu-model-series=<example-model>--restful-portPort exposed by the inference service
--readiness-probe-actionConnection type for readiness probes. Valid values:
httpGet,exec,grpc,tcpSocket--readiness-probe-action-optionConnection method for readiness probes
--readiness-probe-optionReadiness probe configuration
--imageContainer image for the inference service
--dataMounts a PVC into the container. Format:
<pvc-name>:<mount-path>. Runarena data listto list available PVCsTo check the deployment status, run:
arena serve get deepseek-r1Expected output:
Name: deepseek-r1 Namespace: default Type: Custom Version: v1 Desired: 1 Available: 1 Age: 6h Address: 10.0.78.27 Port: RESTFUL:8000 GPU: 1 Instances: NAME STATUS AGE READY RESTARTS GPU NODE ---- ------ --- ----- -------- --- ---- deepseek-r1-v1-custom-serving-54d579d994-dqwxz Running 1h 1/1 0 1 virtual-kubelet-cn-hangzhou-b
Step 3: Verify the inference service
Use
kubectl port-forwardto set up port forwarding from local machine to the inference service.NoteThe port forwarding established by
kubectl port-forwarddoes not provide production-grade reliability, security, or scalability. Therefore, it is intended only for development and debugging purposes and is not suitable for production environments. For more information about production-ready networking solutions in Kubernetes clusters, see Ingress management.kubectl port-forward svc/deepseek-r1-v1 8000:8000Expected output:
Forwarding from 127.0.0.1:8000 -> 8000 Forwarding from [::1]:8000 -> 8000Send a test request to the OpenAI-compatible chat completions endpoint.
curl http://localhost:8000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "deepseek-r1", "messages": [ { "role": "user", "content": "Write a letter to my daughter from the future 2035 and tell her to study science and technology well, be the master of science and technology, and promote the development of science and technology and economy. She is now in grade 3." } ], "max_tokens": 1024, "temperature": 0.7, "top_p": 0.9, "seed": 10 }'Expected output:
{ "id": "chatcmpl-53613fd815da46df92cc9b92cd156146", "object": "chat.completion", "created": 1739261570, "model": "deepseek-r1", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "<think>...</think>\n\nDear Future 2035: ..." }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 40, "completion_tokens": 994, "total_tokens": 1034 } }
(Optional) Step 4: Clean up
Delete the inference service and storage resources when they are no longer needed.
Delete the inference service.
arena serve delete deepseek-r1Expected output:
INFO[0007] The serving job deepseek-r1 with version v1 has been deleted successfullyDelete the PVC and PV.
kubectl delete pvc llm-model kubectl delete pv llm-modelExpected output:
persistentvolumeclaim "llm-model" deleted persistentvolume "llm-model" deleted
What's next
ACS GPU compute power is also available in Container Service for Kubernetes (ACK) Pro clusters. See Use the computing power of ACS in ACK Pro clusters.
For information about deploying DeepSeek models in ACK, see the following topics:
For detailed information about DeepSeek R1/V3 models, see the following topics:
ACS AI container images are specialized images for ACS clusters using GPU instances. For information about currently available images, see ACS AI container image release notes.