Customize GPU drivers on nodes with an OSS URL
ACK clusters install different default NVIDIA driver versions depending on cluster type and version. To install a higher driver version on GPU nodes, upload a custom driver to OSS and configure node pool labels to pull it by OSS URL.
Usage notes
ACK does not guarantee compatibility between GPU driver versions and CUDA library versions. You are responsible for verifying their compatibility.
For detailed driver requirements for different NVIDIA GPU models, see the official NVIDIA documentation.
If you use a custom OS image with a pre-installed GPU driver, NVIDIA Container Runtime, or other GPU components, ACK cannot guarantee that the custom GPU driver is compatible with other ACK GPU components, such as monitoring agents.
When you specify a GPU driver version by using a node pool label, the driver is installed only on new nodes added to the pool. Existing nodes are not affected. To apply the new driver to existing nodes, you must Remove a node from a cluster or node pooland then Add existing nodes.
The gn7, GPU-accelerated compute-optimized instance familyand GPU compute-optimized: gn, ebm, and sccinstance types have compatibility issues with driver versions 510.xxx and 515.xxx. Use a driver version earlier than 510 with GPU System Processor (GSP) disabled (for example, 470.xxx.xxxx) or version 525.125.06 or later.
ECS instances of the GPU compute-optimized: gn, ebm, and sccor Ebmgn7e instance familyinstance type support only NVIDIA driver versions 525.125.06 or later.
If you customize the GPU driver version by specifying a version number or by using an OSS URL, the OS and driver may become incompatible after the update. See Supported NVIDIA driver versions in ACK to select a compatible driver.
Uploading your own GPU driver to OSS may cause incompatibilities with the OS image, ECS instance type, or container runtime, leading to node creation failure. ACK does not guarantee success with this method. Verify the configuration before proceeding.
Step 1: Download the target driver
If the NVIDIA driver versions supported by ACK do not include your required version, download the driver from the official NVIDIA website. This example uses version 550.90.07. Download the NVIDIA-Linux-x86_64-550.90.07.run file to your local machine.
Step 2: Download NVIDIA Fabric Manager
Download NVIDIA Fabric Manager from the NVIDIA YUM repository. The Fabric Manager version must match the driver version.
wget https://developer.download.nvidia.cn/compute/cuda/repos/rhel7/x86_64/nvidia-fabric-manager-550.90.07-1.x86_64.rpmStep 3: Create an OSS bucket
Log on to the OSS console and create a bucket.
Create the bucket in the same region as your ACK cluster so nodes can pull the driver over the internal network.
Step 4: Upload files to the OSS bucket
Log on to the OSS console and upload the
NVIDIA-Linux-x86_64-550.90.07.runandnvidia-fabric-manager-550.90.07-1.x86_64.rpmfiles to the root directory of the bucket.ImportantUpload files to the root directory of the bucket, not a subdirectory.
On the bucket page, in the left navigation pane, click . In the Actions column for the uploaded file, click Details.
In the Details panel, turn off the Use HTTPS switch.
ImportantACK pulls driver files over HTTP, but OSS defaults to HTTPS. Turn off Use HTTPS to enable HTTP access.
In the left navigation pane, click Overview. Copy the internal endpoint from the lower part of the page.
ImportantExternal endpoints are slow and may cause GPU node creation to fail. Use an internal endpoint (contains
-internal) or an accelerated endpoint (containsoss-accelerate).If a file download fails, adjust the bucket's access control policy.
Step 5: Configure node pool labels
-
Log on to the ACK console. In the left navigation pane, click Clusters.
-
On the Clusters page, click the name of your cluster. In the left navigation pane, click .
Click Create Node Pool to add GPU nodes. See Create and manage a node pool for parameter details. The key parameters for this configuration:
In the Node Labels section, click the
icon to add the following labels. Replace example values with your actual values.Key
Value
ack.aliyun.com/nvidia-driver-oss-endpointInternal endpoint of the OSS bucket from Step 4.
my-nvidia-driver.oss-cn-beijing-internal.aliyuncs.comack.aliyun.com/nvidia-driver-runfileNVIDIA driver filename from Step 1.
NVIDIA-Linux-x86_64-550.90.07.runack.aliyun.com/nvidia-fabricmanager-rpmFabric Manager filename from Step 2.
nvidia-fabric-manager-550.90.07-1.x86_64.rpm
Step 6: Verify the driver installation
View Pods with the
component: nvidia-device-pluginlabel:kubectl get po -n kube-system -l component=nvidia-device-plugin -o wideExpected output:
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES nvidia-device-plugin-cn-beijing.192.168.1.127 1/1 Running 0 6d 192.168.1.127 cn-beijing.192.168.1.127 <none> <none> nvidia-device-plugin-cn-beijing.192.168.1.128 1/1 Running 0 17m 192.168.1.128 cn-beijing.192.168.1.128 <none> <none> nvidia-device-plugin-cn-beijing.192.168.8.12 1/1 Running 0 9d 192.168.8.12 cn-beijing.192.168.8.12 <none> <none> nvidia-device-plugin-cn-beijing.192.168.8.13 1/1 Running 0 9d 192.168.8.13 cn-beijing.192.168.8.13 <none> <none>The output shows that the newly added node's Pod is
nvidia-device-plugin-cn-beijing.192.168.1.128.Verify that the correct driver version is installed:
kubectl exec -ti nvidia-device-plugin-cn-beijing.192.168.1.128 -n kube-system -- nvidia-smiExpected output:
+-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 550.90.07 Driver Version: 550.90.07 CUDA Version: 12.4 | |-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 Tesla P100-PCIE-16GB On | 00000000:00:08.0 Off | Off | | N/A 31C P0 26W / 250W | 0MiB / 16384MiB | 0% Default | | | | N/A | +-----------------------------------------+------------------------+----------------------+ +-----------------------------------------------------------------------------------------+ | Processes: | | GPU GI CI PID Type Process name GPU Memory | | ID ID Usage | |=========================================================================================| | No running processes found | +-----------------------------------------------------------------------------------------+The output shows driver version 550.90.07, confirming the custom NVIDIA driver was installed successfully.
Other methods
You can also set the custom driver OSS URL when creating a node pool with the CreateClusterNodePool API:
{
// Other sections are omitted.
......
"tags": [
{
"key": "ack.aliyun.com/nvidia-driver-oss-endpoint",
"value": "my-nvidia-driver.oss-cn-beijing-internal.aliyuncs.com"
},
{
"key": "ack.aliyun.com/nvidia-driver-runfile",
"value": "NVIDIA-Linux-x86_64-550.90.07.run"
},
{
"key": "ack.aliyun.com/nvidia-fabricmanager-rpm",
"value": "nvidia-fabric-manager-550.90.07-1.x86_64.rpm"
}
],
// Other sections are omitted.
......
}