Container Service for Kubernetes (ACK) provides GPU sharing and scheduling capabilities for model inference scenarios that share a single GPU. It also enforces GPU memory isolation in the kernel mode of the NVIDIA driver. If you have installed the GPU sharing component in your cluster, upgrade the component to the latest version if the GPU driver or operating system on a node is incompatible with the installed cGPU version. This topic describes how to manage the GPU sharing component on GPU nodes to enable GPU scheduling and isolation.
Prerequisites
Before using GPU sharing, activate Cloud-native AI Suite. For billing details, see Billing of Cloud-native AI Suite.
For notes on using cGPU, see cGPU FAQ.
An ACK managed cluster is created, and a GPU cloud server is selected as the instance type.
Limitations
The cGPU memory isolation feature is supported only on Alibaba Cloud Elastic Compute Service (ECS) nodes. Do not add the
ack.node.gpu.schedule=cgpuorack.node.gpu.schedule=core_memlabel to nodes that are not Alibaba Cloud ECS nodes.Do not set the CPU policy to
staticfor nodes that use GPU sharing.If you need to specify a custom path for the KubeConfig file, run the
export KUBECONFIG=<kubeconfig>command. Thekubectl inspect cgpucommand does not support the--kubeconfigparameter.Because cGPU isolation does not support Unified Virtual Memory (UVM), you cannot call cudaMallocManaged() to allocate GPU memory. Instead, call cudaMalloc(). For more information, see the NVIDIA documentation.
The GPU sharing DaemonSet pods do not have the highest priority on a node, so other higher-priority pods may preempt their resources and cause eviction. To prevent this, modify the DaemonSet you are using, such as
gpushare-device-plugin-dsfor shared GPU memory, and addpriorityClassName: system-node-criticalto ensure high priority.For performance, a single GPU card supports a maximum of 20 pods when using cGPU. If this limit is exceeded, subsequent pods scheduled to that card will fail to run and return the error message
Error occurs when creating cGPU instance: unknown.To prevent the scheduler from continuing to schedule cGPU pods to nodes where a single card has already reached the 20-pod limit, run the following command to add a label to the node:
kubectl label nodes <NODE> ack.node.gpu.max-pod=20The value
20is fixed and cannot be changed to other values.You can install the GPU sharing component in any region. However, the GPU memory isolation feature is available only in the following regions. Make sure your cluster is in one of these regions.
Version compatibility.
Item
Supported versions
Kubernetes version
ack-ai-installer versions earlier than 1.12.0 support clusters running Kubernetes 1.18.8 or later.
ack-ai-installer 1.12.0 or later supports only clusters running Kubernetes 1.20 or later.
NVIDIA driver version
418.87.01 or later
Container runtime version
Docker: 19.03.5 or later
containerd: 1.4.3 or later
Operating system
Alibaba Cloud Linux 3.x (container-optimized edition requires ack-ai-installer v1.12.6 or later), Alibaba Cloud Linux 2.x, CentOS 7.6, CentOS 7.7, CentOS 7.9, Ubuntu 22.04
Supported GPUs
P-series, T-series, V-series, A-series, H-series
Install the GPU sharing component
Step 1: Install the GPU sharing component
Cloud-native AI Suite not deployed
-
Log on to the ACK console. In the left navigation pane, click Clusters.
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
On the Cloud-native AI Suite page, click Deploy.
On the Deploy Cloud-native AI Suite page, select Scheduling policies extension (Batch scheduling, GPU sharing, and GPU topology-aware scheduling).
(Optional) To customize the cGPU policy, click Advanced to the right of Scheduling Policy Extension (Batch Task Scheduling, GPU Sharing, Topology-aware GPU Scheduling). In the Parameters dialog box, modify the
policyfield for cGPU and click OK.If you have no special requirements for cGPU compute power sharing, the recommended setting is the default
policy: 5(native scheduling). For more information about cGPU policies, see Install and use the cGPU service.At the bottom of the Cloud-native AI Suite page, click Deploy Cloud-native AI Suite.
After installation, the ack-ai-installer component appears in the component list on the Cloud-native AI Suite page.
Cloud-native AI Suite deployed
-
Log on to the ACK console. In the left navigation pane, click Clusters.
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
Find the ack-ai-installer component and click Deploy in the Actions column.
(Optional) In the Parameters dialog box that appears, modify the
policyfield for cGPU.If you have no special requirements for cGPU compute power sharing, the recommended setting is the default
policy: 5(native scheduling). For more information about cGPU policies, see Install and use the cGPU service.After you finish modifying the parameters, click OK.
After the component is installed, the Status of ack-ai-installer changes to Deployed.
Step 2: Enable GPU sharing and isolation
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
On the Node Pools page, click Create Node Pool. Configure the parameters as described in Create and manage a node pool.
On the Create Node Pool page, configure the parameters for the node pool and then click Confirm. The following table describes the key parameters.
Parameter
Description
Expected number of nodes
Enter the initial number of nodes for the node pool. Use 0 if no nodes are needed initially.
Node labels
Set the label value based on your business requirements. For more information about node labels, see Enable scheduling.
This example uses the label value
cgpu. This value enables GPU sharing on the node, where each pod only needs to request GPU memory resources. Multiple pods on a single GPU card share compute power and have isolated memory.Click
next to Node Labels. Set the Key to ack.node.gpu.scheduleand the Value tocgpu.ImportantFor notes on using the cGPU isolation feature, see cGPU FAQ.
After adding the GPU sharing label, do not use the
kubectl label nodescommand to change the GPU scheduling label or use the label management feature on the Nodes page to change the node label. This can cause potential issues. For more information, see Enable scheduling. The recommended method is described in Enable scheduling.
Step 3: Add GPU nodes
If you already created GPU nodes when you added the node pool, you can skip this step.
After creating the node pool, add GPU nodes to it. When doing so, ensure you specify a GPU cloud server as the instance type. For more information, see Add existing nodes to a node pool or Create and manage a node pool.
Step 4: Install the GPU query tool
Download the kubectl-inspect-cgpu executable to a directory in your PATH, such as
/usr/local/bin/.If you are using Linux, run the following command to download kubectl-inspect-cgpu.
wget http://aliacs-k8s-cn-beijing.oss-cn-beijing.aliyuncs.com/gpushare/kubectl-inspect-cgpu-linux -O /usr/local/bin/kubectl-inspect-cgpuIf you are using macOS, run the following command to download kubectl-inspect-cgpu.
wget http://aliacs-k8s-cn-beijing.oss-cn-beijing.aliyuncs.com/gpushare/kubectl-inspect-cgpu-darwin -O /usr/local/bin/kubectl-inspect-cgpu
Make the file executable:
chmod +x /usr/local/bin/kubectl-inspect-cgpuCheck the GPU usage in the cluster:
kubectl inspect cgpuExpected output:
NAME IPADDRESS GPU0(Allocated/Total) GPU Memory(GiB) cn-shanghai.192.168.6.104 192.168.6.104 0/15 0/15 ---------------------------------------------------------------------- Allocated/Total GPU Memory In Cluster: 0/15 (0%)
Upgrade the GPU sharing component
Step 1: Determine the upgrade method
The upgrade method depends on how the ack-ai-installer component was installed. There are two installation methods:
Installed via the Cloud-native AI Suite (Recommended): The component was installed from the Cloud-native AI Suite page.
Installed from the App Catalog (no longer available): The ack-ai-installer component was installed from the App Catalog page in the Marketplace. This method is now closed, but existing components installed this way can still be upgraded through it.
ImportantIf you uninstall a component that was installed using this method, you must activate the Cloud-native AI Suite and complete the installation to reinstall it.
Step 2: Upgrade the component
Cloud-native AI suite
-
Log on to the ACK console. In the left navigation pane, click Clusters.
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
In the Components section, locate the ack-ai-installer component and click Upgrade in the Actions column.
App catalog
-
Log on to the ACK console. In the left navigation pane, click Clusters.
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
On the Helm page, locate the ack-ai-installer release. In the Actions column, click Update. Follow the prompts to select the latest Chart version and update the component.
ImportantIf you need to customize the Chart configuration, confirm the component update after making your changes.
After the update, check the Helm page to confirm that the Chart version of the ack-ai-installer release is the latest version.
Step 3: Upgrade existing nodes
Upgrading the ack-ai-installer component does not automatically upgrade the cGPU version on existing nodes. Use the following information to determine whether the cGPU isolation feature is enabled on your nodes.
If your cluster contains GPU nodes with the cGPU isolation feature enabled, you must also upgrade the cGPU version on those nodes. For more information, see Upgrade the cGPU version of a node.
If your cluster does not have any nodes with the cGPU isolation feature enabled, skip this step.
NoteIf a node has the label
ack.node.gpu.schedule=cgpuorack.node.gpu.schedule=core_mem, the cGPU isolation feature is enabled.Upgrading the cGPU version on existing nodes requires stopping all business pods on those nodes. Perform this operation during off-peak hours to minimize business disruption.