Associate a CDH computing resource

Updated at:

Associate a Cloudera Distribution for Hadoop (CDH) cluster with your DataWorks workspace as a computing resource for data synchronization and development.

Prerequisites

  • The RAM user must be a member of the workspace and have the workspace administrator role.

  • A CDH cluster is deployed.

    Note

    DataWorks supports CDH clusters deployed outside Alibaba Cloud ECS, provided the deployment environment connects to an Alibaba Cloud VPC. You can typically use the network connectivity method for on-premises data sources to establish a stable connection.

  • A resource group is associated with the workspace, and network connectivity is verified.

Limitations

  • Regions: This feature is available in China (Beijing), China (Shanghai), China (Shenzhen), China (Hangzhou), China (Zhangjiakou), China (Chengdu), and Germany (Frankfurt).

  • Permissions:

    Operator

    Required permissions

    Alibaba Cloud account

    No additional permissions are required.

    RAM user/RAM role

    • Only workspace members with the O&M role, workspace administrator role, or the AliyunDataWorksFullAccess permission can create a computing resource. For more information about how to grant permissions, see Grant users workspace administrator permissions.

Go to the computing resource list page

  1. Log on to the DataWorks console. In the navigation pane on the left, switch to the target region and click More > Management Center. Select your workspace from the drop-down list and click Go to Management Center.

  2. In the navigation pane on the left, click Computing Resources to open the computing resource list page.

Associate a CDH computing resource

On the computing resources page, configure and associate a CDH computing resource.

  1. Select the computing resource type.

    1. Click Associate Computing Resources to go to the Associate Computing Resources page.

    2. On the Associate Computing Resources page, select CDH as the computing resource type. The Associate CDH Computing Resource configuration page appears.

  2. Configure the CDH computing resource.

    On the Associate CDH Computing Resource page, configure the following parameters.

    Parameter

    Description

    Cluster Version

    Select the version of the cluster that you want to register.

    DataWorks has built-in support for CDH 5.16.2, CDH 6.1.1, CDH 6.2.1, CDH 6.3.2, and CDP 7.1.7, with all component versions in the cluster connection information fixed. If none of these versions meet your requirements, select Custom Version and configure the component versions as needed.

    Note
    • The required components vary by cluster version and are displayed on the configuration page.

    • If you register a cluster with a Custom Version, only legacy exclusive resource groups for scheduling are supported. After registration, you must submit a ticket to request that technical support initialize the environment.

    Cluster Name

    Select the name of a cluster registered in another workspace to load its configuration, or specify a custom name to create a new configuration.

    Cluster Connection Information

    Hive connection information

    Used to submit Hive jobs to the cluster.

    • HiveServer2 configuration format: jdbc:hive2://<host>:<port>/<database>

    • Metastore configuration format: thrift://<host>:<port>

    How to obtain parameters: For more information, see Obtain CDH or CDP cluster information and configure network connectivity.

    Component version selection: The system automatically detects the component versions for the current cluster.

    Note

    If you use a serverless resource group to access CDH components by domain name, you must configure authoritative resolution for the domain names and set their effective scope in PrivateZone of Alibaba Cloud DNS.

    Impala connection information

    Used to submit Impala jobs.

    Configuration format: jdbc:impala://<host>:<port>/<schema>.

    Spark connection information

    Select a default Spark version and configure the connection.

    Yarn connection information

    Configurations for submitting tasks and viewing task details.

    • Yarn.Resourcemanager.Address configuration format: http://<host>:<port>

      Note

      The submission address for Spark or MapReduce tasks.

    • Jobhistory.Webapp.Address configuration format: http://<host>:<port2>

      Note

      The web UI address of the JobHistory Server for viewing details of completed tasks.

    MapReduce connection information

    Select a default MapReduce version and configure the connection.

    Presto connection information

    Used to submit Presto jobs.

    JDBC address information configuration format: jdbc:presto://<host>:<port>/<catalog>/<schema>

    Note

    This is not a default CDH component. Configure this component as needed.

    Cluster Configuration File

    Configure core-site file

    Contains global configurations for the Hadoop Core library, such as common I/O settings for HDFS and MapReduce.

    Upload this file to run Spark or MapReduce tasks.

    Configure hdfs-site file

    Contains HDFS-related configurations, such as data block size, number of replicas, and path names.

    Configure mapred-site file

    Configures MapReduce parameters such as execution mode and scheduling behavior.

    Upload this file to run MapReduce tasks.

    Configure yarn-site file

    Contains all configurations related to YARN daemons, such as environment settings for the resource manager, node managers, and application runtime.

    Upload this file if you plan to run Spark or MapReduce tasks or use Kerberos for account mapping.

    Configure hive-site file

    Configures Hive parameters such as database connection information, Hive Metastore settings, and execution engine.

    Upload this file if you select Kerberos as the account mapping type.

    Configure spark-defaults file

    Specifies default Spark job configurations such as memory size and CPU cores. Spark applications apply these settings at runtime.

    Upload this file to run Spark tasks.

    Configure config.properties file

    Contains configurations related to the Presto server, such as global properties for the coordinator and worker nodes in the Presto cluster.

    Upload this file if you use the Presto component and select OPEN LDAP or Kerberos as the account mapping type.

    Configure presto.jks file

    Stores security certificates (private keys and public key certificates) for SSL/TLS encrypted communication between Presto processes.

    Default Access Identity

    To use an identity associated with a mapped cluster account, go to the Computing Resources page, and on the Account Mappings tab, set the cluster identity mapping.

    • Development environment: You can use a cluster account or the mapped cluster account of the task executor.

    • Production environment: You can use a cluster account, the mapped cluster account of the task owner, the mapped cluster account of the Alibaba Cloud account, or the mapped cluster account of a RAM user.

    Computing Resource Instance Name

    Specify a custom name for the computing resource instance. At runtime, you can select a computing resource for a task based on this name.

  3. Click Confirm to complete the configuration.

Initialize the resource group

When you register a cluster for the first time or change cluster service configurations (such as core-site.xml), initialize the resource group to ensure it can access the CDH cluster after you configure network connectivity.

  1. On the Computing Resources page, find the CDH computing resource that you created. In the upper-right corner, click Initialize Resource Group.

  2. Click Initialize next to the resource group that you want to use. After the resource group is initialized, click Determine.

(Optional) Configure YARN resource queue

On the Computing Resources page, find your associated CDH cluster. On the YARN Resource Queue tab, click Edit YARN Resource Queue to set a dedicated YARN resource queue for tasks in different modules.

(Optional) Configure Spark parameters

You can set dedicated Spark parameters for tasks in different modules.

  1. On the Computing Resources page, find your associated CDH cluster.

  2. On the Spark-related Parameter tab, click Edit Spark Parameters to go to the page where you can edit Spark parameters for the CDH cluster.

  3. You can configure the Spark property information by clicking the Add button below the module and entering the Spark Property Name and Spark Property Value.

(Optional) Configure host settings

Task submission may fail when a DataWorks serverless resource group connects to a Kerberos-enabled CDH cluster.

This happens because Kerberos relies on hostnames for authentication. If DNS cannot resolve a cluster IP address to its Kerberos-registered hostname, authentication fails.

The host configuration feature lets you define a static IP-to-hostname mapping for a CDH computing resource. DataWorks prioritizes this mapping when accessing the cluster, ensuring successful Kerberos authentication.

  1. Find the CDH computing resource that you want to configure and click Host Configuration.

  2. In the dialog box that appears, enter the mapping relationship in the IP address hostname format. Enter only one mapping per line.

  3. Click OK to save the configuration.

  4. After saving, the configured hostname information appears on the computing resource card, confirming that the configuration is active.

Important
  • Format requirement: The IP address and hostname must be separated by one or more spaces.

  • Configuration integrity: Configure correct mappings for all critical nodes involved in Kerberos authentication and task execution, such as NameNodes, ResourceManagers, and NodeManagers.

  • Scope: This host configuration applies only to the current computing resource, not to other resources in the workspace.

Next steps

After you configure the CDH computing resource, use CDH-related nodes in Data Studio for data development.