Associate a Kubernetes computing resource
Associate your Kubernetes cluster as a computing resource in DataWorks so you can develop and schedule Spark on Kubernetes tasks in Data Studio.
Usage notes
-
Workspace requirement: Only workspaces with the new Data Studio enabled are supported.
-
Permission requirement:
Operator
Required permissions
Alibaba Cloud account
No additional permissions are required.
RAM user/RAM role
To create a computing resource, you must be a workspace member with the O&M or Workspace Administrator role, or have the
AliyunDataWorksFullAccesspermission. For more information, see Grant a user the permissions of a workspace administrator.
Prerequisites
-
Cluster preparation: Have an active Kubernetes cluster, such as a Container Service for Kubernetes (ACK) cluster, and its kubeconfig file. For more information, see Obtain the kubeconfig file of a cluster.
-
Network connectivity: Associate a serverless resource group with your workspace and ensure it has network connectivity to the Kubernetes cluster's API server.
-
If the cluster is in an Alibaba Cloud VPC : Follow the instructions in Step 2: Establish a network connection to connect the two VPCs.
-
If the cluster is self-managed in an on-premises IDC: Follow the instructions in Connect to an on-premises data source to connect the on-premises network to the VPC.
-
Associate the Kubernetes computing resource
Go to the computing resources page
-
Log on to the DataWorks console . Switch to the target region, and then in the left-side navigation pane, click Workspace .
-
On the workspace list, find the target workspace and click Details in the Actions column to go to the workspace configuration page. From the left-side navigation pane, select Computing Resource.
Associate the Kubernetes computing resource
Configure and associate the Kubernetes computing resource.
-
Select the computing resource type.
-
Click Associate Computing Resource to go to the Associate Computing Resource page.
-
On the Associate Computing Resource page, select Kubernetes as the computing resource type, and the Bind Kubernetes Computing Resource configuration page opens.
-
-
Configure the Kubernetes computing resource.
On the Associate Kubernetes computing resource configuration page, configure the parameters as described in the following table.
Parameter
Description
kubeconfig file
Upload the kubeconfig file to Object Storage Service (OSS) and enter the full path of the file in OSS. The path must follow one of these formats:
-
Root directory example:
oss://your-bucket/.dataworks/kubeconfig -
Subdirectory example:
oss://your-bucket/any/path/you/like/.dataworks/kubeconfig
ImportantYou must authorize the AliyunServiceRoleForDataworksEngine service-linked role. For security reasons, this role can access only the */.dataworks folder. This folder can be in the root directory of the bucket or in any subdirectory.
Dashboard URL
(Optional) The dashboard URL of your Kubernetes cluster. After you configure this URL, you can navigate to the cluster console directly from DataWorks.
Spark Web UI URL
(Optional) The web UI URLs for Spark jobs. After you configure these URLs, you can click the links in job run logs to open the Spark Web UI for real-time monitoring or historical analysis.
-
Spark Web UI URL: The URL of the real-time UI for running Spark jobs. This URL typically points to a unified Spark UI proxy service.
-
Spark History Server URL: The URL of the history UI and logs for completed Spark jobs. This URL typically points to a Spark History Server that you deployed.
Spark History Server URL
Computing resource instance name
Enter a custom name for this Kubernetes computing resource.
Description
(Optional) A description of the computing resource for easier identification and management.
-
-
Test connectivity.
In the Connection Configuration section, select the exclusive resource group that accesses this Kubernetes cluster, and then click Test Connectivity. If the test fails, check the network configurations of the resource group and the cluster.
-
Click Confirm to associate the Kubernetes computing resource.
Next steps
After the Kubernetes computing resource is associated, you can create a Kubernetes Spark task in DataStudio and select it as the compute engine for data development.