Create a CDH Spark node

Updated at:

A CDH Spark node runs a Spark task that you compile in CDH and schedules that task in DataWorks. Create a CDH Spark node to submit Spark jobs to a CDH cluster that is registered with your workspace and to run those jobs on a recurring schedule.

Use cases

DataWorks schedules and monitors the Spark jobs that you submit to a CDH cluster, which simplifies job O&M and makes resource management more efficient. Spark supports complex in-memory analysis and large-scale, low-latency data analytics applications. CDH Spark nodes are commonly used in the following scenarios:

  • Data analysis — Use Spark SQL, Dataset, and the DataFrame API to perform complex data aggregation, filtering, and transformation, and to quickly gain insights into your data.

  • Stream processing — Use Spark Streaming to process real-time data streams and perform immediate analysis and decision-making.

  • Machine learning — Use Spark MLlib for data pre-processing, feature engineering, model training, and evaluation.

  • Large-scale ETL — Extract, transform, and load large datasets to prepare data for data warehouses or other storage systems.

How it works

A CDH Spark task spans two systems. In CDH, you develop and compile the Spark task code into a task JAR file. In DataWorks, you upload the JAR file as a CDH JAR resource, reference the resource from a CDH Spark node, and submit the task to the CDH cluster by using the spark-submit command. DataWorks then runs the node on the scheduling properties that you configure for it.

The last action in the process depends on the mode of your workspace. If your workspace is in standard mode, committing the task is not the final action: the task must also be deployed to the production environment.

Prerequisites

  • A workflow is created in DataStudio.

    In DataStudio, development tasks are organized into workflows. You must create a workflow before you can create a node. For more information, see Create a workflow.

  • A CDH cluster is created and registered with your DataWorks workspace.

    You must register your CDH cluster with a DataWorks workspace before creating CDH nodes and tasks. For more information, see Bind a CDH compute resource in the old version of DataStudio.

  • A compiled task JAR file is available.

    Before you use DataWorks to schedule a CDH Spark task, develop and compile the Spark task code in CDH to generate the compiled task JAR file. For guidance on developing a CDH Spark task, see Spark Overview.

  • (Optional) If you are using a RAM user, the user must be added to the workspace and assigned the Development or Workspace Administrator role. The Workspace Administrator role has extensive permissions, so assign it with caution. For more information on adding members, see Add members to a workspace.

  • (Recommended) A serverless resource group is purchased and configured.

    The configuration includes binding the resource group to your workspace and setting up the network. For more information, see Use a serverless resource group. For the other resource group type that CDH Spark tasks support, see Limitations.

Limitations

Confirm the following limits before you create a CDH Spark node:

  • Resource groups — A CDH Spark task runs on a serverless resource group (recommended) or an old-version exclusive resource group for scheduling.

  • Network connectivity — If the task needs to access the public internet or a VPC, you must select a scheduling resource group with the necessary network connectivity. For more information, see Network connectivity solutions.

  • Resource file size — A JAR file that you upload as a CDH JAR resource cannot exceed 50 MB.

  • Kerberos authentication — If Kerberos authentication is enabled, grant the current user write permissions on the storage path directory before you upload a resource.

  • Comments — The CDH Spark node editor does not support comments. If the node code contains comments, an error occurs when you run the node.

Step 1: Create a CDH Spark node

  1. Log on to the DataWorks console. In the target region, click Data Development and O&M > Data Development in the left-side navigation pane. Select a workspace from the drop-down list and click Go to Data Development.

  2. Right-click a workflow and choose Create Node > cdh > CDH Spark.

  3. In the Create Node dialog box, configure the engine instance, path, name, and other information for the node.

  4. Click OK to create the node. You can then develop and configure the corresponding task in the node that you created.

Step 2: Create a CDH JAR resource and write the node code

In the CDH Spark node that you created, reference a JAR resource, write the node code, and submit the task by using the spark-submit command:

  1. Create a CDH JAR resource.

    In the corresponding workflow, right-click cdh > Resources, choose Create Resource > CDH JAR, and in the Create Resource dialog box, click Click Upload and select the file that you want to upload.

    Configure the Storage Path (/user/admin/lib by default) and enter a Name for the resource. If the resource type is JAR, the file name must include the .jar extension. For the size limit of the resource file and for the permissions that Kerberos authentication requires, see Limitations.

  2. Reference the CDH JAR resource.

    1. Open the CDH node that you created and stay on the editing page.

    2. In cdh > Resources, find the resource that you want to reference, such as spark-examples_2.11-2.4.0.jar, right-click the resource name, and select Insert Resource Path.

      After you select Insert Resource Path, if a statement in the ##@resource_reference{""} format appears on the code editing page of the CDH node, the code resource is referenced. The following example shows the inserted statement and the name of the referenced resource:

    ##@resource_reference{"spark-examples_2.11-2.4.0.jar"}
    spark-examples_2.11-2.4.0.jar
  3. Modify the CDH Spark node code to add the spark-submit command.

    Important

    Rewrite your task code based on the following example and do not add comments. For more information, see Limitations.

    The following example shows the modified code:

    ##@resource_reference{"spark-examples_2.11-2.4.0.jar"}
    spark-submit --class org.apache.spark.examples.SparkPi --master yarn  spark-examples_2.11-2.4.0.jar 100
    • org.apache.spark.examples.SparkPi: The main class of the task in the JAR file that you compiled.

    • spark-examples_2.11-2.4.0.jar: The name of the CDH JAR resource that you uploaded.

Step 3: Configure task scheduling

If you need to run the task on a recurring schedule, click Scheduling in the right-side pane to configure its scheduling properties:

Step 4: Debug the code

  1. (Optional) Select a runtime resource group and assign values to custom parameters.

  2. Save and run the SQL statements.

    In the toolbar, click the 保存 icon to save the SQL statements, and then click the 运行 icon to run the task.

  3. Save and run the SQL statements.

    In the toolbar, click the 保存 icon to save the SQL statements, and then click the 运行 icon to run the task.

  4. (Optional) Perform smoke testing.

    To run smoke testing in the development environment, you can do so during the commit process or after you commit the node. For more information, see Perform smoke testing.

Next steps

  1. Commit and deploy the node task.

    1. Click the icon in the toolbar to save the node.

    2. Click the icon in the toolbar to commit the node task.

    3. In the Commit Node dialog box, enter a Change Description.

    4. Click Determine.

      If you use a workspace in standard mode, click Deploy on the left side of the top menu bar to deploy the task to the production environment after the task is committed. For more information, see Deploy tasks.

  2. View the periodic task.

    1. In the upper-right corner of the editing page, click O&M Personnel to go to Operation Center in the production environment.

    2. View the periodic tasks that are running. For more information, see Manage periodic tasks.

      If you want to view more details about periodic tasks, click Operation Center in the top menu bar. For more information, see Operation Center overview.