Configuring scheduling dependencies

Updated at:

Scheduling dependencies define the upstream-downstream relationships between periodically scheduled nodes in DataWorks. After you configure a scheduling dependency, the system triggers a downstream node instance only after all its upstream instances succeed, ensuring that data is produced and consumed in the correct order.

How it works

A scheduling dependency specifies that a node starts only after its upstream nodes succeed. Once configured, the DataWorks scheduling system automatically orchestrates the execution order. A downstream instance is triggered only when all its upstream instances have succeeded and all other conditions, such as time and resource availability, are met.

DataWorks establishes dependencies by matching the output names of upstream nodes with the input names of downstream nodes. The core workflow for configuring a dependency is as follows:

  1. Configure the output on the upstream node: Add an output name to the upstream node, usually in the format project_name.table_name (for example, my_project.dim_user), to represent the data table produced by the node.

  2. Configure the input on the downstream node: In the downstream node, search for and select the output name of the upstream node as its input (dependency). This establishes the dependency relationship.

  3. Automatic parsing (optional): For SQL-based nodes, DataWorks can automatically parse INSERT and SELECT statements in your code to identify input and output tables and then generate the dependency configuration. You can also manually adjust the automatically parsed configuration. For a list of node types that support automatic parsing, see Automatic parsing scenarios for different node types.

Important

Each node must have at least one output name. The system automatically generates a default output for each node in the format project_name.nodeID_out. This default output is retained even if you delete all custom outputs.

Rules and limitations

  • Effective upon deployment: A scheduling dependency takes effect only after you submit and deploy the node to Operation Center. Configurations made in the development environment are not automatically synchronized to the production environment.

  • Upstream and downstream scheduling status: For a dependency to work, both the upstream and downstream node instances must be generated and in a normal scheduling state. If a node is misconfigured or an upstream instance is abnormal, the node may become isolated and cannot be scheduled normally.

  • Circular dependency restriction: The system prohibits circular dependencies (A depends on B, and B depends on A), including both direct and indirect cycles. If a circular dependency is detected during submission, the system blocks the deployment and returns an error.

Dependency types

DataWorks supports two types of scheduling dependencies: same-cycle dependency and cross-cycle dependency, each applicable to different business scenarios.

Prerequisite concepts

A cycle is defined by a node's schedule settings. It refers to the time offset between two adjacent scheduled instances of a node, determined by its scheduling frequency. For example, for a daily scheduled task, the previous cycle is the instance from the previous day. For an hourly scheduled task, the previous cycle is the instance from the previous hour.

Scheduling frequency

One cycle

Daily, weekly, monthly, or yearly schedule

1 day

Note

For weekly, monthly, or yearly scheduled tasks, instances are still generated on a daily basis (instances on non-scheduled days are dry run instances). Therefore, dependency calculations are based on the day granularity, and the previous cycle instance may be in a dry run state.

Hourly schedule

Hourly interval

Minute-level schedule

Minute-level interval (for example, every 5 minutes)

Two dependency types

Example: A daily scheduled node A produces table dim_user, and downstream node B consumes this table:

  • Same-cycle dependency: The instance of B for today waits until the instance of A for today succeeds before running. That is, B consumes data produced by A on the same day.

  • Cross-cycle dependency: The instance of B for today waits until the instance of A for yesterday succeeds before running. That is, B consumes data produced by A on the previous day.

Comparison item

Same-cycle dependency

Cross-cycle dependency (depends on previous cycle)

Meaning

The current cycle instance of this node depends on the result of the upstream node's instance in the same cycle.

The current cycle instance of this node depends on the result of a specified node's instance in the previous cycle. The specified node can be this node itself (self-dependency), a direct downstream child node, or any other node.

Representation in DAG

Displayed as a solid line.

Displayed as a dashed line.

Typical scenarios

Node B needs to read data produced by node A today.

A node depends on data produced the previous day (such as T-1 data retrieval); hourly/minute-level tasks use self-dependency to achieve serial execution and avoid concurrent execution across multiple cycles.

Configuration method

Supports automatic parsing, workflow drag-and-drop connection, and manual addition.

In the "Previous Cycle" section of the schedule settings panel, select the dependency form and specify the node ID.

Note: Same-cycle dependency and cross-cycle dependency can coexist between the same pair of nodes, but you must clearly define the business meaning of each. If you only need a cross-cycle dependency, remember to delete the same-cycle dependency automatically generated by the system. Otherwise, the downstream instance still waits for the upstream instance in the current cycle to complete before running, resulting in unexpected delays.

Scheduling dependency configuration guide

All nodes must have upstream dependencies configured before they can be deployed to Operation Center for automatic scheduling. If no data dependency exists, the node must depend on a virtual node or root node. When configuring dependencies, review the business logic, clarify the dependency targets and types, and select the most appropriate configuration method to build a robust and structurally clear data workflow.

1. Identify dependency targets

Before configuring dependencies, complete the following preparations:

  • Review lineage: Confirm whether the tables/partitions produced by the upstream node match the tables/partitions read by the downstream node.

  • Check schedule settings: Ensure that the node's schedule, effective time, scheduling parameters, and other settings are correctly configured, because schedule settings directly affect dependency behavior.

Select the dependency target based on how the current node depends on data.

Scenario 1: Depends on the direct output of an upstream node

  • Applicable scenario: The data required by the downstream node comes directly from a table produced by another upstream node that is automatically scheduled by DataWorks.

  • Configuration strategy: We strongly recommend that you configure node dependencies based on data lineage.

  • Core value: This is the most direct and robust approach. The scheduling system ensures that the downstream node always starts only after the upstream data is ready, guaranteeing end-to-end data consistency.

Scenario 2: Depends on non-scheduled upstream data (data-readiness-driven)

  • Applicable scenario: The upstream data is not managed by the DataWorks scheduling system and cannot generate scheduled instances for downstream dependencies. Examples include:

    • Files pushed to OSS/FTP by external business systems;

    • Tables produced by real-time synchronization;

    • Tables produced by third-party synchronization tools (not scheduled by DataWorks);

    • Temporary tables manually uploaded or produced by manual execution.

  • Configuration strategy: Configure a check node (such as a Check node) to actively verify whether data is ready (for example, check whether a file exists or whether a table partition has been generated). The downstream business node then depends on this check node.

  • Core value: This converts "data production" into a "scheduling event", enabling data-readiness-driven downstream processes and ensuring data correctness in non-scheduling-chain scenarios.

Scenario 3: No direct data dependency, but business logic association exists

  • Applicable scenario: The node is completely independent in terms of data processing/code logic, but from a business logic perspective, it needs to belong to a workflow or be scheduled periodically.

  • Configuration strategy:

    • Depend on a virtual node: You can aggregate a group of related tasks for unified management, forming a logical unit that facilitates unified start/stop, monitoring, and maintenance, keeping business logic well-organized.

    • Depend on the workspace root node: This ensures that the task is properly instantiated by the scheduling system and executed on time, preventing it from becoming an isolated node that cannot be automatically scheduled.

  • Core value: Prevents node isolation, makes workflow start/stop and status monitoring clearer, and ensures the completeness of business logic.

2. Select the dependency type

If the current node depends on the direct output of an upstream node (Scenario 1), further determine whether the dependent data is from the same cycle or a previous cycle of the upstream node.

Core judgment

Determine which cycle's output the downstream node actually reads from the upstream node. In most scenarios, the data written by a node to a specific partition is dynamically determined by scheduling parameters. See Scheduling parameters to understand how scheduling parameters are replaced. If you need to depend on a node within the same workspace, check its scheduling parameter configuration.

How to confirm

  • Same-workspace node: Check the scheduling parameters in the upstream node's code. Determine whether the partition written after parameter replacement is for "today" or "yesterday".

    • In the development environment, check the upstream node's scheduling parameter configuration and code details. In the production environment, check the parameter replacement results in the instance details.

  • Cross-workspace node: Use Data Map to view the upstream table's partition information and change history.

    • Confirm the actual partition value written each day.

Select the type

  • The downstream code reads the upstream partition for the current day/current cycle: Same-cycle dependency.

  • The downstream code reads the upstream partition for the previous day/previous cycle: Cross-cycle dependency.

  • Hourly/minute-level tasks need serial execution without concurrency: Cross-cycle dependency (self-dependency).

Important

Consequences of not correctly confirming lineage:

  1. Missing dependency risk: If a table lineage exists but no scheduling dependency is configured, the downstream task starts before the upstream instance succeeds, resulting in missing data or incomplete data.

  2. Parameter mismatch risk: If a dependency is configured but partition parameters are misaligned (for example, the upstream produces today's partition but the downstream reads yesterday's partition), this results in data logic errors and quality anomalies.

3. Configure the dependency

Based on the dependency targets and types confirmed in steps 1 and 2, select the appropriate configuration method to configure the dependency.

DataWorks supports dependencies between tasks with different scheduling frequencies. Combined with same-cycle/cross-cycle dependencies and scheduling parameters, this enables a wide variety of scheduling scenarios. For details, see:

4. Verify the scheduling dependency

After configuration and before deployment, you must verify the dependency:

Verification method

Description

During configuration: Preview dependencies

Preview whether the current scheduling dependency configuration meets expectations before deployment.

  • Currently, only the immediate upstream and downstream dependencies of the current node can be viewed.

  • To ensure that the current task's dependencies are correct, confirm that the upstream nodes are in a saved state.

  • In the dependency preview diagram, solid lines represent same-cycle dependencies, and dashed lines represent cross-cycle dependencies (dependencies on the previous cycle).

During submission: Compare code parsing results

Confirm whether the dependency changes in the current version meet expectations and assess the impact on production during node submission.

When automatic parsing is enabled, to ensure normal production data output, you need to perform a secondary confirmation of scheduling changes when submitting a node. Use this feature to ensure that dependency changes do not affect production task data output.

After deployment: View scheduled tasks

Confirm in Operation Center that the dependencies of the production scheduled task meet expectations after node deployment.

  • Confirm the scheduling dependencies of production tasks

    In a standard mode workspace, the node dependencies in the development environment and production environment can be different. The scheduling dependencies for production environment nodes must be configured in Data Studio and take effect only after deployment.

    After the node is deployed, you can go to the Scheduled Tasks page in Operation Center, expand the upstream and downstream of the current task, and view the scheduling dependency status.

    Important

    The Scheduled Tasks page displays the latest state of nodes in the production environment. However, whether a scheduled instance has newly added or removed dependencies is related to the selected instance generation method.

  • Confirm the data status of production tasks

    After confirming the scheduling dependencies, you also need to verify the partition data read and write behavior of upstream and downstream nodes (that is, whether the scheduling parameter configuration is correct). This prevents the downstream node from encountering data quality issues because the upstream node produces data that is not what the current node depends on.

    Note

    If the task deployment process involves workflow controls, we recommend that after deployment, you go to the Scheduled Tasks page in Operation Center to check the task's scheduling dependencies and related properties. If you find that the task does not meet expectations, verify whether the deployment process is blocked. For details, see Deploy nodes.

Impact of removing dependencies

During task iteration, you may need to remove or adjust existing scheduling dependencies.

Before removing a dependency, evaluate the impact on downstream task scheduling behavior to avoid task isolation or data incidents.

Downstream dependency scenario

Impact after removal

Risk level

Downstream depends only on the current node

The downstream task becomes an isolated node, loses its upstream trigger mechanism, and is no longer automatically scheduled.

High

Downstream depends on multiple parent nodes

The downstream task may start before upstream data is ready, resulting in missing data or computation errors.

Medium

Downstream depends on a cross-cycle instance

If the cross-cycle dependency is removed, the downstream node may read data from the wrong business date, resulting in data logic confusion.

Medium

Use cases

  • Offline data warehouse layered construction: Configure end-to-end dependencies across the ODS → DWD → DWS → ADS pipeline to ensure that data in each layer is produced in order.

  • Standard ETL pipeline: Configure same-cycle dependencies to ensure that downstream tasks strictly wait for upstream instances to succeed before execution, guaranteeing the order and consistency of the data processing pipeline.

  • Next-day (T+1) reports: Configure cross-cycle dependencies (offset -1) so that today's task depends on yesterday's complete business data, enabling accurate next-day data analysis and output.

  • Multi-frequency mixed aggregation: Configure cross-cycle dependencies so that daily-granularity tasks depend on all cycle instances of hourly-granularity tasks, ensuring that underlying data is fully ready before aggregation.

  • External data readiness trigger: Configure custom dependencies or check nodes to confirm that external files have arrived or interfaces are ready before triggering the workflow, enabling cross-system scheduling coordination.

  • Complex workflow control: Use virtual nodes to aggregate multi-branch dependencies as workflow control milestones, simplifying pipeline structure and improving monitoring visibility.

FAQ

The following are typical scenarios. For more frequently asked questions about scheduling dependencies, see FAQ about dependency relationships.

  • Node uniqueness.

    • Nodes have different forms in the development and production environments but remain unique: The scheduling dependency configuration of the same node can differ between the development environment and the production environment. That is, the same node can have two different forms in the development and production environments, but the node is unique.

    • Remove downstream dependencies in both environments before undeploying a node: Due to node uniqueness, to ensure that downstream tasks retrieve data and run correctly, before undeploying an upstream task in DataWorks, you must first remove the dependency in the downstream node's schedule settings, then reconfigure the upstream nodes that the downstream node needs to depend on, and submit and deploy the changes. The upstream task can only be undeployed after the dependency is removed in both the development and production environments.

  • Related to instance generation method.

    • When creating nodes, ensure that the upstream and downstream nodes use the same instance generation method. Otherwise, the difference in instance generation methods may cause the upstream node to generate instances on the current day while the downstream node generates instances the next day, resulting in downstream instances becoming isolated instances.

    • If you change the schedule of an existing node and select the option to generate instances immediately after deployment, the already generated instances are not automatically deleted when you modify the node's scheduling dependencies. The dependencies of the day's scheduled instances may become inconsistent after node deployment. For details, see Impact of immediate instance generation on the current day's scheduled instance dependencies.

  • Error indicating more than 200 upstream dependencies when updating a job by using OpenAPI.

    • Error details: 'One file could not have more than 200 inputs 'One file could not have more than 200 inputs'.

    • You can add virtual nodes between the upstream and downstream in Data Studio to reduce the number of direct upstream dependencies for the current node. For details about how to configure virtual nodes, see Create a virtual node