Associate a Kubernetes computing resource

Updated at:

Associate your Kubernetes cluster as a computing resource in DataWorks so you can develop and schedule Spark on Kubernetes tasks in Data Studio.

Usage notes

  • Workspace requirement: Only workspaces with the new Data Studio enabled are supported.

  • Permission requirement:

    Operator

    Required permissions

    Alibaba Cloud account

    No additional permissions are required.

    RAM user/RAM role

    To create a computing resource, you must be a workspace member with the O&M or Workspace Administrator role, or have the AliyunDataWorksFullAccess permission. For more information, see Grant a user the permissions of a workspace administrator.

Prerequisites

  • Cluster preparation: Have an active Kubernetes cluster, such as a Container Service for Kubernetes (ACK) cluster, and its kubeconfig file. For more information, see Obtain the kubeconfig file of a cluster.

  • Network connectivity: Associate a serverless resource group with your workspace and ensure it has network connectivity to the Kubernetes cluster's API server.

Associate the Kubernetes computing resource

Go to the computing resources page

  1. Log on to the DataWorks console . Switch to the target region, and then in the left-side navigation pane, click Workspace .

  2. On the workspace list, find the target workspace and click Details in the Actions column to go to the workspace configuration page. From the left-side navigation pane, select Computing Resource.

Associate the Kubernetes computing resource

Configure and associate the Kubernetes computing resource.

  1. Select the computing resource type.

    1. Click Associate Computing Resource to go to the Associate Computing Resource page.

    2. On the Associate Computing Resource page, select Kubernetes as the computing resource type, and the Bind Kubernetes Computing Resource configuration page opens.

  2. Configure the Kubernetes computing resource.

    On the Associate Kubernetes computing resource configuration page, configure the parameters as described in the following table.

    Parameter

    Description

    kubeconfig file

    Upload the kubeconfig file to Object Storage Service (OSS) and enter the full path of the file in OSS. The path must follow one of these formats:

    • Root directory example: oss://your-bucket/.dataworks/kubeconfig

    • Subdirectory example: oss://your-bucket/any/path/you/like/.dataworks/kubeconfig

    Important

    You must authorize the AliyunServiceRoleForDataworksEngine service-linked role. For security reasons, this role can access only the */.dataworks folder. This folder can be in the root directory of the bucket or in any subdirectory.

    Dashboard URL

    (Optional) The dashboard URL of your Kubernetes cluster. After you configure this URL, you can navigate to the cluster console directly from DataWorks.

    Spark Web UI URL

    (Optional) The web UI URLs for Spark jobs. After you configure these URLs, you can click the links in job run logs to open the Spark Web UI for real-time monitoring or historical analysis.

    • Spark Web UI URL: The URL of the real-time UI for running Spark jobs. This URL typically points to a unified Spark UI proxy service.

    • Spark History Server URL: The URL of the history UI and logs for completed Spark jobs. This URL typically points to a Spark History Server that you deployed.

    Spark History Server URL

    Computing resource instance name

    Enter a custom name for this Kubernetes computing resource.

    Description

    (Optional) A description of the computing resource for easier identification and management.

  3. Test connectivity.

    In the Connection Configuration section, select the exclusive resource group that accesses this Kubernetes cluster, and then click Test Connectivity. If the test fails, check the network configurations of the resource group and the cluster.

  4. Click Confirm to associate the Kubernetes computing resource.

Next steps

After the Kubernetes computing resource is associated, you can create a Kubernetes Spark task in DataStudio and select it as the compute engine for data development.