ACS supports multiple GPU families. Use the alibabacloud.com/gpu-model-series label in your cluster to specify a GPU family. Different GPU families are suited to different use cases. You can select an instance family based on your specific needs.
GU8TF
A compute GPU.
Features 96 GB of memory per GPU and native support for the FP8 data format. This is ideal for single-node inference on models with 70 billion or more parameters.
Provides high-speed NVLink interconnects between all eight GPUs, suitable for small- to medium-scale model training. A 1.6 Tbps Remote Direct Memory Access (RDMA) interconnect accelerates node-to-node communication.
The following table specifies the available resource configurations for Pods using this compute GPU.
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
1 (96 GB) | 2 | 2–16 | 1 | 30–256 |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
12 | 12–96 | 1 | ||
14 | 14–112 | 1 | ||
16 | 16–128 | 1 | ||
22 | 22, 32, 64, 128 | N/A | ||
2 (96 GB) | 16 | 16–128 | 1 | 30–512 |
32 | 32, 64, 128, 230 | N/A | ||
46 | 64, 128, 230 | N/A | ||
4 (96 GB) | 32 | 32, 64, 128, 256 | N/A | 30–1,024 |
64 | 64, 128, 256, 460 | N/A | ||
92 | 128, 256, 460 | N/A | ||
8 (96 GB) | 64 | 64, 128, 256, 512 | N/A | 30–2,048 |
128 | 128, 256, 512, 920 | N/A | ||
184 | 256, 512, 920 | N/A |
GU8TEF
A compute GPU.
Features 141 GB of GPU memory and supports the FP8 floating-point format. Multi-GPU configurations support single-node inference for DeepSeek R1.
NVLink interconnects eight GPUs for small- to medium-scale model training. A 1.6 Tbps high-speed RDMA interconnect accelerates node-to-node communication.
The Pod specifications for this compute GPU are:
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
1 (141 GB) | 2 | 2–16 | 1 | 30–768 |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
12 | 12–96 | 1 | ||
14 | 14–112 | 1 | ||
16 | 16–128 | 1 | ||
22 | 22, 32, 64, 128, 225 | N/A | ||
2 (141 GB) | 16 | 16–128 | 1 | 30–1,536 |
32 | 32, 64, 128, 256 | N/A | ||
46 | 64, 128, 256, 450 | N/A | ||
4 (141 GB) | 32 | 32, 64, 128, 256 | N/A | 30–3,072 |
64 | 64, 128, 256, 512 | N/A | ||
92 | 128, 256, 512, 900 | N/A | ||
8 (141 GB) | 64 | 64, 128, 256, 512 | N/A | 30–6,144 |
128 | 128, 256, 512, 1,024 | N/A | ||
184 | 256, 512, 1,024, 1,800 | N/A |
L20 (GN8IS)
This family provides compute GPUs.
Supports common acceleration features such as TensorRT, the FP8 floating-point format, and peer-to-peer (P2P) communication between GPUs.
Provides 48 GB of GPU memory per GPU, enabling single-node inference on models with 70 billion or more parameters in multi-GPU configurations.
The following table lists the available resource configurations for Pods in this family.
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
1 (48 GB) | 2 | 2–16 | 1 | 30–256 |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
12 | 12–96 | 1 | ||
14 | 14–112 | 1 | ||
16 | 16–120 | 1 | ||
2 (48 GB) | 16 | 16–128 | 1 | 30–512 |
32 | 32, 64, 128, 230 | N/A | ||
4 (48 GB) | 32 | 32, 64, 128, 256 | N/A | 30–1,024 |
64 | 64, 128, 256, 460 | N/A | ||
8 (48 GB) | 64 | 64, 128, 256, 512 | N/A | 30–2,048 |
128 | 128, 256, 512, 920 | N/A |
L20X (GX8SF)
A compute GPU.
Offers 141 GB of memory per GPU, enabling single-node inference for large models in multi-GPU configurations.
Features NVLink interconnects between all eight GPUs, ideal for large-scale model training and inference. A 3.2 Tbps high-speed RDMA interconnect accelerates node-to-node communication.
This table lists the Pod specifications for this GPU.
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
8 (141 GB) | 184 | 1,800 | N/A | 30–6,144 |
P16EN
The P16EN is a GPU compute card.
Provides 96 GB of video memory per GPU and supports the FP16 floating-point format. Multi-GPU configurations enable single-node inference for models such as DeepSeek R1.
Provides a 700 GB/s high-speed interconnect between all 16 GPUs, ideal for small- to medium-scale model training. A 1.6 Tbps high-speed RDMA interconnect accelerates node-to-node communication.
The Pod specification constraints for this GPU card are as follows:
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage |
1 (96 GB) | 2 | 2–16 | 1 | 30–384 GB |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
2 (96 GB) | 4 | 4–32 | 1 | 30–768 GB |
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
16 | 16–128 | 1 | ||
22 | 32, 64, 128, 225 | N/A | ||
4 (96 GB) | 8 | 8–64 | 1 | 30–1536 GB |
16 | 16–128 | 1 | ||
32 | 32, 64, 128, 256 | N/A | ||
46 | 64, 128, 256, 450 | N/A | ||
8 (96 GB) | 16 | 16–128 | 1 | 30–3072 GB |
32 | 32, 64, 128, 256 | N/A | ||
64 | 64, 128, 256, 512 | N/A | ||
92 | 128, 256, 512, 900 | N/A | ||
16 (96 GB) | 32 | 32, 64, 128, 256 | N/A | 30–6144 GB |
64 | 64, 128, 256, 512 | N/A | ||
128 | 128, 256, 512, 1,024 | N/A | ||
184 | 256, 512, 1,024, 1,800 | N/A |
T4
The T4 is a GPU compute card.
Built on the Turing architecture, each GPU provides 16 GB of GPU memory and 320 GB/s of memory bandwidth.
The variable-precision Tensor Cores deliver up to 65 TFLOPS (FP16), 130 TOPS (INT8), and 260 TOPS (INT4).
This GPU has the following Pod resource constraints:
Instance family | GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
single-node family | 1 (16 GB GPU memory) | 2 | 2–8 | 1 | 30–1536 |
4 | 4–16 | 1 | |||
6 | 6–24 | 1 | |||
8 | 8–32 | 1 | |||
10 | 10–40 | 1 | |||
12 | 12–48 | 1 | |||
14 | 14–56 | 1 | |||
16 | 16–64 | 1 | |||
24 | 24, 48, 90 | N/A | |||
2 (16 GB GPU memory) | 16 | 16–64 | 1 | ||
24 | 24, 48, 96 | N/A | |||
32 | 32, 64, 128 | N/A | |||
48 | 48, 96, 180 | N/A |
A10
A GPU.
Based on the Ampere architecture, each GPU has 24 GB of memory and supports common acceleration features such as RTX and TensorRT.
Pod resource constraints for this GPU:
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
1 (24 GB) | 2 | 2–8 | 1 | 30–256 |
4 | 4–16 | 1 | ||
6 | 6–24 | 1 | ||
8 | 8–32 | 1 | ||
10 | 10–40 | 1 | ||
12 | 12–48 | 1 | ||
14 | 14–56 | 1 | ||
16 | 16–60 | 1 | ||
2 (24 GB x 2) | 16 | 16–64 | N/A | 30–512 |
32 | 32, 64, 120 | N/A | ||
4 (24 GB x 4) | 32 | 32, 64, 128 | N/A | 30–1,024 |
64 | 64, 128, 240 | N/A | ||
8 (24 GB x 8) | 64 | 64, 128, 256 | N/A | 30–2,048 |
128 | 128, 256, 480 | N/A |
G28Ti
GPU card.
Provides 11 GB of gpu memory per GPU and supports common acceleration features such as TensorRT and CUDA. It also supports NVLink and P2P communication between GPUs.
This table lists the pod specifications for this GPU card.
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
1 (11 GB gpu memory) | 2 | 2–8 | 1 | 30–1,536 |
4 | 4–16 | 1 | ||
6 | 6–24 | 1 | ||
8 | 8–32 | 1 | ||
10 | 10–40 | 1 | ||
12 | 12–48 | 1 |
G49E
The G49E is a GPU compute card.
Each GPU provides 48 GB of GPU memory, supports common acceleration features such as RTX and TensorRT, and enables peer-to-peer (P2P) communication between GPUs.
This table details the pod specifications for this GPU card.
GPU | vCPU | Memory (GiB) | Increment (GiB) | Storage (GiB) |
1 (48 GB) | 2 | 2–16 | 1 | 30–256 |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
12 | 12–96 | 1 | ||
14 | 14–112 | 1 | ||
16 | 16–120 | 1 | ||
2 (48 GB x 2) | 16 | 16–128 | 1 | 30–512 |
32 | 32, 64, 128, 230 | N/A | ||
4 (48 GB x 4) | 32 | 32, 64, 128, 256 | N/A | 30–1,024 |
64 | 64, 128, 256, 460 | N/A | ||
8 (48 GB x 8) | 64 | 64, 128, 256, 512 | N/A | 30–2,048 |
128 | 128, 256, 512, 920 | N/A |
G59
GPU.
Each GPU has 32 GB of GPU memory, supports acceleration features like RTX and TensorRT, and enables peer-to-peer (P2P) communication between GPUs.
The following table lists the Pod specification constraints for this GPU.
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
1 (32 GB GPU memory) | 2 | 2–16 | 1 | 30–256 |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
12 | 12–96 | 1 | ||
14 | 14–112 | 1 | ||
16 | 16–128 | 1 | ||
22 | 22, 32, 64, 128 | N/A | ||
2 (32 GB GPU memory each) | 16 | 16–128 | 1 | 30–512 |
32 | 32, 64, 128, 256 | N/A | ||
46 | 64, 128, 256, 360 | N/A | ||
4 (32 GB GPU memory each) | 32 | 32, 64, 128, 256 | N/A | 30–1,024 |
64 | 64, 128, 256, 512 | N/A | ||
92 | 128, 256, 512, 720 | N/A | ||
8 (32 GB GPU memory each) | 64 | 64, 128, 256, 512 | N/A | 30–2,048 |
128 | 128, 256, 512, 1,024 | N/A | ||
184 | 256, 512, 1,024, 1,440 | N/A |
L20N
The L20N is a GPU compute card.
Features the new Blackwell architecture, high-frequency CPUs, and large memory.
Ideal for cost-effective GPU acceleration in various scenarios, including training for autonomous driving and embodied intelligence, large model inference, film and animation rendering, metaverse applications, and cloud gaming.
The specifications for Pods that support this GPU card are as follows:
GPU | vCPU | Memory (GiB) | Memory step size (GiB) | Storage (GiB) |
1 (48 GB VRAM) | 2 | 2–16 | 1 | 30–2,048 |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
12 | 12–96 | 1 | ||
14 | 14–112 | 1 | ||
16 | 16–128 | 1 | ||
32 | 32, 64, 128, 256 | N/A | ||
2 (2 × 48 GB VRAM) | 16 | 16–128 | 1 | |
32 | 32, 64, 128, 256 | N/A | ||
64 | 64, 128, 256, 512 | N/A | ||
4 (4 × 48 GB VRAM) | 32 | 32, 64, 128, 256 | N/A | |
64 | 64, 128, 256, 512 | N/A | ||
128 | 128, 256, 512, 1,024 | N/A | ||
8 (8 × 48 GB VRAM) | 64 | 64, 128, 256, 512 | N/A | |
128 | 128, 256, 512, 1,024 | N/A | ||
256 | 256, 512, 1,024, 2,048 | N/A |
L20NE
GPU card.
Provides a high-frequency CPU, large-capacity memory, and a professional graphics card based on the new Blackwell architecture.
Delivers a cost-effective GPU cloud service for GPU-accelerated workloads such as autonomous driving, embodied intelligence training, large model inference, film and animation rendering, metaverse, and cloud gaming.
The following are the available Pod specifications:
GPU | vCPU | Memory (GiB) | Memory increment (GiB) | Storage (GiB) |
1 (72 GB GPU memory) | 2 | 2–16 | 1 | 30–2,048 |
4 | 4–32 | 1 | ||
6 | 6–48 | 1 | ||
8 | 8–64 | 1 | ||
10 | 10–80 | 1 | ||
12 | 12–96 | 1 | ||
14 | 14–112 | 1 | ||
16 | 16–128 | 1 | ||
32 | 32, 64, 128, 256 | N/A | ||
2 (72 GB GPU memory each) | 16 | 16–128 | 1 | |
32 | 32, 64, 128, 256 | N/A | ||
64 | 64, 128, 256, 512 | N/A | ||
4 (72 GB GPU memory each) | 32 | 32, 64, 128, 256 | N/A | |
64 | 64, 128, 256, 512 | N/A | ||
128 | 128, 256, 512, 1,024 | N/A | ||
8 (72 GB GPU memory each) | 64 | 64, 128, 256, 512 | N/A | |
128 | 128, 256, 512, 1,024 | N/A | ||
256 | 256, 512, 1,024, 2,048 | N/A |