Customize GPU drivers on nodes with an OSS URL

Updated at:

ACK clusters install different default NVIDIA driver versions depending on cluster type and version. To install a higher driver version on GPU nodes, upload a custom driver to OSS and configure node pool labels to pull it by OSS URL.

Usage notes

Step 1: Download the target driver

If the NVIDIA driver versions supported by ACK do not include your required version, download the driver from the official NVIDIA website. This example uses version 550.90.07. Download the NVIDIA-Linux-x86_64-550.90.07.run file to your local machine.

Step 2: Download NVIDIA Fabric Manager

Download NVIDIA Fabric Manager from the NVIDIA YUM repository. The Fabric Manager version must match the driver version.

wget https://developer.download.nvidia.cn/compute/cuda/repos/rhel7/x86_64/nvidia-fabric-manager-550.90.07-1.x86_64.rpm

Step 3: Create an OSS bucket

Log on to the OSS console and create a bucket.

Note

Create the bucket in the same region as your ACK cluster so nodes can pull the driver over the internal network.

Step 4: Upload files to the OSS bucket

  1. Log on to the OSS console and upload the NVIDIA-Linux-x86_64-550.90.07.run and nvidia-fabric-manager-550.90.07-1.x86_64.rpm files to the root directory of the bucket.

    Important

    Upload files to the root directory of the bucket, not a subdirectory.

  2. On the bucket page, in the left navigation pane, click File Management > Objects. In the Actions column for the uploaded file, click Details.

  3. In the Details panel, turn off the Use HTTPS switch.

    Important

    ACK pulls driver files over HTTP, but OSS defaults to HTTPS. Turn off Use HTTPS to enable HTTP access.

  4. In the left navigation pane, click Overview. Copy the internal endpoint from the lower part of the page.

    Important
    • External endpoints are slow and may cause GPU node creation to fail. Use an internal endpoint (contains -internal) or an accelerated endpoint (contains oss-accelerate).

    • If a file download fails, adjust the bucket's access control policy.

Step 5: Configure node pool labels

  1. Log on to the ACK console. In the left navigation pane, click Clusters.

  2. On the Clusters page, click the name of your cluster. In the left navigation pane, click Nodes > Node Pools.

  3. Click Create Node Pool to add GPU nodes. See Create and manage a node pool for parameter details. The key parameters for this configuration:

    In the Node Labels section, click the 1 icon to add the following labels. Replace example values with your actual values.

    Key

    Value

    ack.aliyun.com/nvidia-driver-oss-endpoint

    Internal endpoint of the OSS bucket from Step 4.

    my-nvidia-driver.oss-cn-beijing-internal.aliyuncs.com

    ack.aliyun.com/nvidia-driver-runfile

    NVIDIA driver filename from Step 1.

    NVIDIA-Linux-x86_64-550.90.07.run

    ack.aliyun.com/nvidia-fabricmanager-rpm

    Fabric Manager filename from Step 2.

    nvidia-fabric-manager-550.90.07-1.x86_64.rpm

Step 6: Verify the driver installation

  1. View Pods with the component: nvidia-device-plugin label:

    kubectl get po -n kube-system -l component=nvidia-device-plugin -o wide

    Expected output:

    NAME                                            READY   STATUS    RESTARTS   AGE   IP              NODE                       NOMINATED NODE   READINESS GATES
    nvidia-device-plugin-cn-beijing.192.168.1.127   1/1     Running   0          6d    192.168.1.127   cn-beijing.192.168.1.127   <none>           <none>
    nvidia-device-plugin-cn-beijing.192.168.1.128   1/1     Running   0          17m   192.168.1.128   cn-beijing.192.168.1.128   <none>           <none>
    nvidia-device-plugin-cn-beijing.192.168.8.12    1/1     Running   0          9d    192.168.8.12    cn-beijing.192.168.8.12    <none>           <none>
    nvidia-device-plugin-cn-beijing.192.168.8.13    1/1     Running   0          9d    192.168.8.13    cn-beijing.192.168.8.13    <none>           <none>

    The output shows that the newly added node's Pod is nvidia-device-plugin-cn-beijing.192.168.1.128.

  2. Verify that the correct driver version is installed:

    kubectl exec -ti nvidia-device-plugin-cn-beijing.192.168.1.128 -n kube-system -- nvidia-smi

    Expected output:

    +-----------------------------------------------------------------------------------------+
    | NVIDIA-SMI 550.90.07              Driver Version: 550.90.07      CUDA Version: 12.4     |
    |-----------------------------------------+------------------------+----------------------+
    | GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
    | Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
    |                                         |                        |               MIG M. |
    |=========================================+========================+======================|
    |   0  Tesla P100-PCIE-16GB           On  |   00000000:00:08.0 Off |                  Off |
    | N/A   31C    P0             26W /  250W |       0MiB /  16384MiB |      0%      Default |
    |                                         |                        |                  N/A |
    +-----------------------------------------+------------------------+----------------------+
                                                                                             
    +-----------------------------------------------------------------------------------------+
    | Processes:                                                                              |
    |  GPU   GI   CI        PID   Type   Process name                              GPU Memory |
    |        ID   ID                                                               Usage      |
    |=========================================================================================|
    |  No running processes found                                                             |
    +-----------------------------------------------------------------------------------------+

    The output shows driver version 550.90.07, confirming the custom NVIDIA driver was installed successfully.

Other methods

You can also set the custom driver OSS URL when creating a node pool with the CreateClusterNodePool API:

{
  // Other sections are omitted.
  ......
    "tags": [
      {
        "key": "ack.aliyun.com/nvidia-driver-oss-endpoint",
        "value": "my-nvidia-driver.oss-cn-beijing-internal.aliyuncs.com"
      },
      {
        "key": "ack.aliyun.com/nvidia-driver-runfile",
        "value": "NVIDIA-Linux-x86_64-550.90.07.run"
      },
      {
        "key": "ack.aliyun.com/nvidia-fabricmanager-rpm",
        "value": "nvidia-fabric-manager-550.90.07-1.x86_64.rpm"
      }
    ],
  // Other sections are omitted.
  ......
}