Create a CDH Spark node
A CDH Spark node runs a Spark task that you compile in CDH and schedules that task in DataWorks. Create a CDH Spark node to submit Spark jobs to a CDH cluster that is registered with your workspace and to run those jobs on a recurring schedule.
Use cases
DataWorks schedules and monitors the Spark jobs that you submit to a CDH cluster, which simplifies job O&M and makes resource management more efficient. Spark supports complex in-memory analysis and large-scale, low-latency data analytics applications. CDH Spark nodes are commonly used in the following scenarios:
Data analysis — Use Spark SQL, Dataset, and the DataFrame API to perform complex data aggregation, filtering, and transformation, and to quickly gain insights into your data.
Stream processing — Use Spark Streaming to process real-time data streams and perform immediate analysis and decision-making.
Machine learning — Use Spark MLlib for data pre-processing, feature engineering, model training, and evaluation.
Large-scale ETL — Extract, transform, and load large datasets to prepare data for data warehouses or other storage systems.
How it works
A CDH Spark task spans two systems. In CDH, you develop and compile the Spark task code into a task JAR file. In DataWorks, you upload the JAR file as a CDH JAR resource, reference the resource from a CDH Spark node, and submit the task to the CDH cluster by using the spark-submit command. DataWorks then runs the node on the scheduling properties that you configure for it.
The last action in the process depends on the mode of your workspace. If your workspace is in standard mode, committing the task is not the final action: the task must also be deployed to the production environment.
Prerequisites
A workflow is created in DataStudio.
In DataStudio, development tasks are organized into workflows. You must create a workflow before you can create a node. For more information, see Create a workflow.
A CDH cluster is created and registered with your DataWorks workspace.
You must register your CDH cluster with a DataWorks workspace before creating CDH nodes and tasks. For more information, see Bind a CDH compute resource in the old version of DataStudio.
A compiled task JAR file is available.
Before you use DataWorks to schedule a CDH Spark task, develop and compile the Spark task code in CDH to generate the compiled task JAR file. For guidance on developing a CDH Spark task, see Spark Overview.
(Optional) If you are using a RAM user, the user must be added to the workspace and assigned the Development or Workspace Administrator role. The Workspace Administrator role has extensive permissions, so assign it with caution. For more information on adding members, see Add members to a workspace.
(Recommended) A serverless resource group is purchased and configured.
The configuration includes binding the resource group to your workspace and setting up the network. For more information, see Use a serverless resource group. For the other resource group type that CDH Spark tasks support, see Limitations.
Limitations
Confirm the following limits before you create a CDH Spark node:
Resource groups — A CDH Spark task runs on a serverless resource group (recommended) or an old-version exclusive resource group for scheduling.
Network connectivity — If the task needs to access the public internet or a VPC, you must select a scheduling resource group with the necessary network connectivity. For more information, see Network connectivity solutions.
Resource file size — A JAR file that you upload as a CDH JAR resource cannot exceed 50 MB.
Kerberos authentication — If Kerberos authentication is enabled, grant the current user write permissions on the storage path directory before you upload a resource.
Comments — The CDH Spark node editor does not support comments. If the node code contains comments, an error occurs when you run the node.
Step 1: Create a CDH Spark node
Log on to the DataWorks console. In the target region, click in the left-side navigation pane. Select a workspace from the drop-down list and click Go to Data Development.
Right-click a workflow and choose .
In the Create Node dialog box, configure the engine instance, path, name, and other information for the node.
Click OK to create the node. You can then develop and configure the corresponding task in the node that you created.
Step 2: Create a CDH JAR resource and write the node code
In the CDH Spark node that you created, reference a JAR resource, write the node code, and submit the task by using the spark-submit command:
Create a CDH JAR resource.
In the corresponding workflow, right-click , choose , and in the Create Resource dialog box, click Click Upload and select the file that you want to upload.
Configure the Storage Path (
/user/admin/libby default) and enter a Name for the resource. If the resource type is JAR, the file name must include the.jarextension. For the size limit of the resource file and for the permissions that Kerberos authentication requires, see Limitations.Reference the CDH JAR resource.
Open the CDH node that you created and stay on the editing page.
In , find the resource that you want to reference, such as
spark-examples_2.11-2.4.0.jar, right-click the resource name, and select Insert Resource Path.After you select Insert Resource Path, if a statement in the
##@resource_reference{""}format appears on the code editing page of the CDH node, the code resource is referenced. The following example shows the inserted statement and the name of the referenced resource:
##@resource_reference{"spark-examples_2.11-2.4.0.jar"} spark-examples_2.11-2.4.0.jarModify the CDH Spark node code to add the
spark-submitcommand.ImportantRewrite your task code based on the following example and do not add comments. For more information, see Limitations.
The following example shows the modified code:
##@resource_reference{"spark-examples_2.11-2.4.0.jar"} spark-submit --class org.apache.spark.examples.SparkPi --master yarn spark-examples_2.11-2.4.0.jar 100org.apache.spark.examples.SparkPi: The main class of the task in the JAR file that you compiled.spark-examples_2.11-2.4.0.jar: The name of the CDH JAR resource that you uploaded.
Step 3: Configure task scheduling
If you need to run the task on a recurring schedule, click Scheduling in the right-side pane to configure its scheduling properties:
Configure the basic scheduling properties. For more information, see Configure basic properties.
Configure the scheduling cycle, rerun properties, and dependencies. For more information, see Configure time properties and Configure same-cycle scheduling dependencies.
NoteYou must configure the Rerun attribute properties and specify the Parent Nodes before committing the node.
Configure resource properties. For more information, see Configure resource properties. The scheduling resource group that you select must provide the network connectivity that the task requires. For more information, see Limitations.
Step 4: Debug the code
(Optional) Select a runtime resource group and assign values to custom parameters.
In the toolbar, click the
icon. In the Parameter dialog box, select the resource group to use for debugging. For the resource group types that a CDH Spark task supports, see Limitations.If your task code uses scheduling parameters, assign their values here for debugging. For more information about the value assignment logic, see What is the difference in value assignment logic between Run, Advanced Run, and development-environment smoke testing?.
-
Save and run the SQL statements.
In the toolbar, click the
icon to save the SQL statements, and then click the
icon to run the task. (Optional) Perform smoke testing.
To run smoke testing in the development environment, you can do so during the commit process or after you commit the node. For more information, see Perform smoke testing.
Save and run the SQL statements.
In the toolbar, click the
icon to save the SQL statements, and then click the
icon to run the task.
Next steps
Commit and deploy the node task.
Click the icon in the toolbar to save the node.
Click the icon in the toolbar to commit the node task.
In the Commit Node dialog box, enter a Change Description.
Click Determine.
If you use a workspace in standard mode, click Deploy on the left side of the top menu bar to deploy the task to the production environment after the task is committed. For more information, see Deploy tasks.
View the periodic task.
In the upper-right corner of the editing page, click O&M Personnel to go to Operation Center in the production environment.
View the periodic tasks that are running. For more information, see Manage periodic tasks.
If you want to view more details about periodic tasks, click Operation Center in the top menu bar. For more information, see Operation Center overview.