Patch a critical TOCTOU container escape vulnerability in NVIDIA Container Toolkit 1.17.7 and earlier on GPU-accelerated nodes.
See NVIDIA Security Bulletin 5659.
Affected scope
Your cluster is affected if both conditions are true:
-
Kubernetes version is 1.32 or earlier
-
At least one GPU-accelerated node runs NVIDIA Container Toolkit 1.17.7 or earlier
Does not affect deployments using the Container Device Interface (CDI).
Known attack scenarios require running a malicious container image and accessing GPU resources through the NVIDIA Container Toolkit.
Check the NVIDIA Container Toolkit version
Log on to each GPU-accelerated node and run:
nvidia-container-cli --version
Sample output (unaffected version):
cli-version: 1.17.8
lib-version: 1.17.8
build date: 2025-05-30T13:47+00:00
build revision: 6eda4d76c8c5f8fc174e4abca83e513fb4dd63b0
build compiler: x86_64-linux-gnu-gcc-7 7.5.0
build platform: x86_64
If cli-version is 1.17.7 or earlier, the node is affected.
Preventive measures
While you prepare to apply the fix, restrict image pulls to trusted registries. Enable the container security policy rule in the policy governance feature.
Solution
New GPU-accelerated nodes
ACK edge clusters running Kubernetes 1.20 or later
Nodes created on or after August 4, 2025 automatically install NVIDIA Container Toolkit 1.17.8. No action is required.
Clusters running Kubernetes earlier than 1.20
Upgrade the cluster before creating new nodes to ensure they receive the patched version.
Existing GPU-accelerated nodes
All GPU-accelerated nodes created before August 4, 2025 require a manual fix.
-
Cloud nodes: Follow Vulnerability CVE-2025-23266.
-
Edge nodes: Follow the procedure below.
Apply fixes in batches to maintain system stability.
Fix edge nodes
The fix drains each node, upgrades NVIDIA Container Toolkit to 1.17.8, and restores it to service.
Prerequisites
Ensure that you have:
-
SSH access to the target GPU-accelerated edge node
-
kubectlaccess to the cluster with sufficient permissions to cordon and drain nodes
Step 1: Drain the node
Draining safely migrates workloads to other nodes before you apply changes. Before you proceed, verify that the selected node is the target edge node.
-
Mark the node as unschedulable:
kubectl cordon <NODE_NAME> -
Drain the node:
kubectl drain <NODE_NAME> --grace-period=120 --ignore-daemonsets=true
Step 2: Apply the fix
Log on to the affected node.
-
Set the
REGIONandINTERCONNECT_MODEenvironment variables. Replace example values with your configuration:Parameter Description Example REGIONRegion ID of your ACK edge cluster. See Supported regions. cn-hangzhouINTERCONNECT_MODENetwork access type: basic(Internet access) orprivate(leased line access).basicexport REGION="cn-hangzhou" INTERCONNECT_MODE="basic" -
Run the fix script:
#!/bin/bash set -e if [[ $REGION == "" ]];then echo "Error: REGION is null" exit 1 fi if [[ $INTERCONNECT_MODE == "" ]]; then echo "Error: INTERCONNECT_MODE is null" exit 1 fi NV_TOOLKIT_VERSION=1.17.8 INTERNAL=$( [ "$INTERCONNECT_MODE" = "private" ] && echo "-internal" || echo "" ) PACKAGE=upgrade_nvidia-container-toolkit-${NV_TOOLKIT_VERSION}.tar.gz cd /tmp export PKG_URL_PREFIX="http://aliacs-k8s-${REGION}.oss-${REGION}${INTERNAL}.aliyuncs.com" curl -o ${PACKAGE} ${PKG_URL_PREFIX}/public/pkg/nvidia-container-runtime/${PACKAGE} tar -xf ${PACKAGE} cd pkg/nvidia-container-runtime/upgrade/common bash upgrade-nvidia-container-toolkit.sh -
Verify the output:
Output Meaning INFO No need to upgrade current nvidia-container-toolkit(1.17.8)Already running the patched version. No changes made. INFO succeed to upgrade nvidia container toolkitNode successfully patched.
Step 3: Restore the node
Return the node to service:
kubectl uncordon <NODE_NAME>
Step 4 (optional): Verify GPU functionality
Deploy a GPU workload to confirm the node works correctly. Use sample YAML templates from:
-
Exclusive GPU: Use the default GPU scheduling mode
-
Shared GPU: GPU sharing examples