You can use an OceanBase node in DataWorks to develop and periodically schedule OceanBase tasks, and integrate them with other types of jobs. This topic describes the main workflow for developing tasks using an OceanBase node.
Background information
OceanBase is a distributed relational database developed by Ant Group and Alibaba. It features strong data consistency, high availability, high performance, online scalability, and low costs. OceanBase is also highly compatible with SQL standards and mainstream relational databases. For more information, see What is OceanBase.
Prerequisites
-
A workflow is created.
You must create a workflow before you create a node. For more information, see Create a workflow.
-
You must add your OceanBase database as an ApsaraDB for OceanBase data source in DataWorks. For more information about how to create a data source, see Data source management. For more information about how to use an ApsaraDB for OceanBase data source in DataWorks, see ApsaraDB for OceanBase data source.
NoteOceanBase nodes support only OceanBase data sources that are created with a JDBC connection string.
A network connection is established between your data source and a resource group.
You must make sure that the desired data source is connected to the resource group that you want to use. For more information about how to configure network connectivity, see Establish a network connection between a resource group and a data source.
-
(Optional) If you use a RAM user for task development, the RAM user must be added to the workspace and granted the Development or Workspace Administrator role. The Workspace Administrator role provides extensive permissions. Grant this role with caution. For more information about how to add members to a workspace and grant roles to the members, see Add members to a workspace.
Limitations
Supported regions: China (Hangzhou), China (Shanghai), China (Beijing), China (Zhangjiakou), China (Shenzhen), China (Chengdu), China (Hong Kong), Singapore, Malaysia (Kuala Lumpur), Germany (Frankfurt), US (Silicon Valley), and US (Virginia).
Step 1: Create an OceanBase node
Go to the DataStudio page.
Log on to the DataWorks console. In the top navigation bar, select the desired region. In the left-side navigation pane, choose . On the page that appears, select the desired workspace from the drop-down list and click Go to Data Development.
-
Right-click the target workflow and choose .
-
In the Create Node dialog box, enter a Name for the node and click OK. After the node is created, you can use it to develop and configure your task.
Step 2: Develop an OceanBase task
(Optional) Select an OceanBase data source
If your workspace has multiple OceanBase data sources, you must select which one to use on the node configuration page. If it has only one, it is selected by default.
OceanBase nodes support only OceanBase data sources that are created with a JDBC connection string.
Write SQL code: Basic example
In the code editor of the OceanBase node, write the SQL code for your task. The following code is an example.
SELECT * FROM usertablename;
Write SQL code: Use scheduling parameters
DataWorks Scheduling Parameter let you dynamically pass values to code in periodically scheduled tasks. You can define variables in your code by using the ${variable_name} format. In the right-side navigation pane, click the Scheduling tab and assign values to your variables in the Scheduling Parameter section. For more information about the supported formats and how to configure scheduling parameters, see Supported formats for scheduling parameters and Configure and use scheduling parameters.
The following code is an example.
SELECT '${var}'; -- Example of using a scheduling parameter.
Step 3: Configure task scheduling
If you need to run the node task periodically, click the scheduling properties tab in the right-side navigation pane of the node editor page and configure scheduling properties for the task based on your business requirements. For more information, see Overview of task scheduling properties.
Before you can commit the node, you must specify the Rerun attribute property and Parent Nodes for the node.
Step 4: Debug the task
To debug the task and check whether it runs as expected:
-
(Optional) Select a resource group for scheduling and specify values for custom parameters.
-
Click the
icon in the toolbar. In the Parameter dialog box, select the resource group for scheduling that you want to use for debugging. -
If your task code uses scheduling parameters, you can specify values for them for debugging. For more information about how to specify parameter values, see Debug a task.
-
-
Save and run the task code.
Click the
icon in the toolbar to save the task code and click the
icon to run the task. -
(Optional) Perform smoke testing.
To check if the scheduled task runs as expected in the development environment, you can perform smoke testing when you commit the node or after it is committed. For more information, see Perform smoke testing.
Step 5: Commit and deploy the task
After you configure the node task, you must commit and deploy it. Once deployed, the node runs on a schedule based on its scheduling properties.
-
Click the
icon in the toolbar to save the node. -
Click the
icon in the toolbar to commit the node task.In the Submission dialog box, enter a Change Description. You can also select whether to perform a code review after you commit the node.
Note-
Before you can commit the node, you must specify the Rerun attribute property and Parent Nodes for the node.
-
Code review helps ensure code quality and prevents task errors that may result from deploying incorrect code to the production environment without being reviewed. If you enable code review, the submitted node code can be deployed only after reviewers approve it. For more information, see Code review.
-
If you use a workspace in standard mode, you must click Deploy in the upper-right corner of the node editor page after you commit the task. This operation deploys the task to the production environment. For more information, see Deploy tasks.
Next steps
Task O&M: After the task is committed and deployed, it runs periodically based on its scheduling configuration. You can click Operation Center in the upper-right corner of the node editor to monitor its scheduling and operational status. For more information, see Manage auto-triggered tasks.