Configuring scheduling dependencies
Scheduling dependencies define the upstream-downstream relationships between periodically scheduled nodes in DataWorks. After you configure a scheduling dependency, the system triggers a downstream node instance only after all its upstream instances succeed, ensuring that data is produced and consumed in the correct order.
How it works
A scheduling dependency specifies that a node starts only after its upstream nodes succeed. Once configured, the DataWorks scheduling system automatically orchestrates the execution order. A downstream instance is triggered only when all its upstream instances have succeeded and all other conditions, such as time and resource availability, are met.
DataWorks establishes dependencies by matching the output names of upstream nodes with the input names of downstream nodes. The core workflow for configuring a dependency is as follows:
-
Configure the output on the upstream node: Add an output name to the upstream node, usually in the format
project_name.table_name(for example,my_project.dim_user), to represent the data table produced by the node. -
Configure the input on the downstream node: In the downstream node, search for and select the output name of the upstream node as its input (dependency). This establishes the dependency relationship.
-
Automatic parsing (optional): For SQL-based nodes, DataWorks can automatically parse
INSERTandSELECTstatements in your code to identify input and output tables and then generate the dependency configuration. You can also manually adjust the automatically parsed configuration. For a list of node types that support automatic parsing, see Automatic parsing scenarios for different node types.
Each node must have at least one output name. The system automatically generates a default output for each node in the format project_name.nodeID_out. This default output is retained even if you delete all custom outputs.
Rules and limitations
-
Effective upon deployment: A scheduling dependency takes effect only after you submit and deploy the node to Operation Center. Configurations made in the development environment are not automatically synchronized to the production environment.
-
Upstream and downstream scheduling status: For a dependency to work, both the upstream and downstream node instances must be generated and in a normal scheduling state. If a node is misconfigured or an upstream instance is abnormal, the node may become isolated and cannot be scheduled normally.
-
Circular dependency restriction: The system prohibits circular dependencies (A depends on B, and B depends on A), including both direct and indirect cycles. If a circular dependency is detected during submission, the system blocks the deployment and returns an error.
Dependency types
DataWorks supports two types of scheduling dependencies: same-cycle dependency and cross-cycle dependency, each applicable to different business scenarios.
Prerequisite concepts
A cycle is defined by a node's schedule settings. It refers to the time offset between two adjacent scheduled instances of a node, determined by its scheduling frequency. For example, for a daily scheduled task, the previous cycle is the instance from the previous day. For an hourly scheduled task, the previous cycle is the instance from the previous hour.
|
Scheduling frequency |
One cycle |
|
Daily, weekly, monthly, or yearly schedule |
1 day Note
For weekly, monthly, or yearly scheduled tasks, instances are still generated on a daily basis (instances on non-scheduled days are dry run instances). Therefore, dependency calculations are based on the day granularity, and the previous cycle instance may be in a dry run state. |
|
Hourly schedule |
Hourly interval |
|
Minute-level schedule |
Minute-level interval (for example, every 5 minutes) |
Two dependency types
Example: A daily scheduled node A produces table dim_user, and downstream node B consumes this table:
-
Same-cycle dependency: The instance of B for today waits until the instance of A for today succeeds before running. That is, B consumes data produced by A on the same day.
-
Cross-cycle dependency: The instance of B for today waits until the instance of A for yesterday succeeds before running. That is, B consumes data produced by A on the previous day.
|
Comparison item |
Same-cycle dependency |
Cross-cycle dependency (depends on previous cycle) |
|
Meaning |
The current cycle instance of this node depends on the result of the upstream node's instance in the same cycle. |
The current cycle instance of this node depends on the result of a specified node's instance in the previous cycle. The specified node can be this node itself (self-dependency), a direct downstream child node, or any other node. |
|
Representation in DAG |
Displayed as a solid line. |
Displayed as a dashed line. |
|
Typical scenarios |
Node B needs to read data produced by node A today. |
A node depends on data produced the previous day (such as T-1 data retrieval); hourly/minute-level tasks use self-dependency to achieve serial execution and avoid concurrent execution across multiple cycles. |
|
Configuration method |
Supports automatic parsing, workflow drag-and-drop connection, and manual addition. |
In the "Previous Cycle" section of the schedule settings panel, select the dependency form and specify the node ID. |
Note: Same-cycle dependency and cross-cycle dependency can coexist between the same pair of nodes, but you must clearly define the business meaning of each. If you only need a cross-cycle dependency, remember to delete the same-cycle dependency automatically generated by the system. Otherwise, the downstream instance still waits for the upstream instance in the current cycle to complete before running, resulting in unexpected delays.
Scheduling dependency configuration guide
All nodes must have upstream dependencies configured before they can be deployed to Operation Center for automatic scheduling. If no data dependency exists, the node must depend on a virtual node or root node. When configuring dependencies, review the business logic, clarify the dependency targets and types, and select the most appropriate configuration method to build a robust and structurally clear data workflow.
1. Identify dependency targets
Before configuring dependencies, complete the following preparations:
-
Review lineage: Confirm whether the tables/partitions produced by the upstream node match the tables/partitions read by the downstream node.
-
Check schedule settings: Ensure that the node's schedule, effective time, scheduling parameters, and other settings are correctly configured, because schedule settings directly affect dependency behavior.
Select the dependency target based on how the current node depends on data.
|
Scenario 1: Depends on the direct output of an upstream node |
|
|
Scenario 2: Depends on non-scheduled upstream data (data-readiness-driven) |
|
|
Scenario 3: No direct data dependency, but business logic association exists |
|
2. Select the dependency type
If the current node depends on the direct output of an upstream node (Scenario 1), further determine whether the dependent data is from the same cycle or a previous cycle of the upstream node.
Core judgment
Determine which cycle's output the downstream node actually reads from the upstream node. In most scenarios, the data written by a node to a specific partition is dynamically determined by scheduling parameters. See Scheduling parameters to understand how scheduling parameters are replaced. If you need to depend on a node within the same workspace, check its scheduling parameter configuration.
How to confirm
-
Same-workspace node: Check the scheduling parameters in the upstream node's code. Determine whether the partition written after parameter replacement is for "today" or "yesterday".
-
In the development environment, check the upstream node's scheduling parameter configuration and code details. In the production environment, check the parameter replacement results in the instance details.
-
-
Cross-workspace node: Use Data Map to view the upstream table's partition information and change history.
-
Confirm the actual partition value written each day.
-
Select the type
-
The downstream code reads the upstream partition for the current day/current cycle: Same-cycle dependency.
-
The downstream code reads the upstream partition for the previous day/previous cycle: Cross-cycle dependency.
-
Hourly/minute-level tasks need serial execution without concurrency: Cross-cycle dependency (self-dependency).
Consequences of not correctly confirming lineage:
-
Missing dependency risk: If a table lineage exists but no scheduling dependency is configured, the downstream task starts before the upstream instance succeeds, resulting in missing data or incomplete data.
-
Parameter mismatch risk: If a dependency is configured but partition parameters are misaligned (for example, the upstream produces today's partition but the downstream reads yesterday's partition), this results in data logic errors and quality anomalies.
3. Configure the dependency
Based on the dependency targets and types confirmed in steps 1 and 2, select the appropriate configuration method to configure the dependency.
DataWorks supports dependencies between tasks with different scheduling frequencies. Combined with same-cycle/cross-cycle dependencies and scheduling parameters, this enables a wide variety of scheduling scenarios. For details, see:
4. Verify the scheduling dependency
After configuration and before deployment, you must verify the dependency:
|
Verification method |
Description |
|
Preview whether the current scheduling dependency configuration meets expectations before deployment.
|
|
|
Confirm whether the dependency changes in the current version meet expectations and assess the impact on production during node submission. When automatic parsing is enabled, to ensure normal production data output, you need to perform a secondary confirmation of scheduling changes when submitting a node. Use this feature to ensure that dependency changes do not affect production task data output. |
|
|
Confirm in Operation Center that the dependencies of the production scheduled task meet expectations after node deployment.
|
Impact of removing dependencies
During task iteration, you may need to remove or adjust existing scheduling dependencies.
Before removing a dependency, evaluate the impact on downstream task scheduling behavior to avoid task isolation or data incidents.
|
Downstream dependency scenario |
Impact after removal |
Risk level |
|
Downstream depends only on the current node |
The downstream task becomes an isolated node, loses its upstream trigger mechanism, and is no longer automatically scheduled. |
High |
|
Downstream depends on multiple parent nodes |
The downstream task may start before upstream data is ready, resulting in missing data or computation errors. |
Medium |
|
Downstream depends on a cross-cycle instance |
If the cross-cycle dependency is removed, the downstream node may read data from the wrong business date, resulting in data logic confusion. |
Medium |
Use cases
-
Offline data warehouse layered construction: Configure end-to-end dependencies across the ODS → DWD → DWS → ADS pipeline to ensure that data in each layer is produced in order.
-
Standard ETL pipeline: Configure same-cycle dependencies to ensure that downstream tasks strictly wait for upstream instances to succeed before execution, guaranteeing the order and consistency of the data processing pipeline.
-
Next-day (T+1) reports: Configure cross-cycle dependencies (offset -1) so that today's task depends on yesterday's complete business data, enabling accurate next-day data analysis and output.
-
Multi-frequency mixed aggregation: Configure cross-cycle dependencies so that daily-granularity tasks depend on all cycle instances of hourly-granularity tasks, ensuring that underlying data is fully ready before aggregation.
-
External data readiness trigger: Configure custom dependencies or check nodes to confirm that external files have arrived or interfaces are ready before triggering the workflow, enabling cross-system scheduling coordination.
-
Complex workflow control: Use virtual nodes to aggregate multi-branch dependencies as workflow control milestones, simplifying pipeline structure and improving monitoring visibility.
FAQ
The following are typical scenarios. For more frequently asked questions about scheduling dependencies, see FAQ about dependency relationships.
-
Node uniqueness.
-
Nodes have different forms in the development and production environments but remain unique: The scheduling dependency configuration of the same node can differ between the development environment and the production environment. That is, the same node can have two different forms in the development and production environments, but the node is unique.
-
Remove downstream dependencies in both environments before undeploying a node: Due to node uniqueness, to ensure that downstream tasks retrieve data and run correctly, before undeploying an upstream task in DataWorks, you must first remove the dependency in the downstream node's schedule settings, then reconfigure the upstream nodes that the downstream node needs to depend on, and submit and deploy the changes. The upstream task can only be undeployed after the dependency is removed in both the development and production environments.
-
-
Related to instance generation method.
-
When creating nodes, ensure that the upstream and downstream nodes use the same instance generation method. Otherwise, the difference in instance generation methods may cause the upstream node to generate instances on the current day while the downstream node generates instances the next day, resulting in downstream instances becoming isolated instances.
-
If you change the schedule of an existing node and select the option to generate instances immediately after deployment, the already generated instances are not automatically deleted when you modify the node's scheduling dependencies. The dependencies of the day's scheduled instances may become inconsistent after node deployment. For details, see Impact of immediate instance generation on the current day's scheduled instance dependencies.
-
-
Error indicating more than 200 upstream dependencies when updating a job by using OpenAPI.
-
Error details: 'One file could not have more than 200 inputs 'One file could not have more than 200 inputs'.
-
You can add virtual nodes between the upstream and downstream in Data Studio to reduce the number of direct upstream dependencies for the current node. For details about how to configure virtual nodes, see Create a virtual node
-