Do-while node

Updated at:

DataWorks provides the do-while node, which implements "execute first, then check" loop logic. This node is ideal when you need to repeat a series of tasks until a dynamic condition is met, such as polling to check for a file's existence or processing data in batches until the source is empty. You can orchestrate a task flow within the loop body and use a dedicated End node to control the loop's exit.

Use cases

In data development, the do-while node simplifies the design of workflows that need to repeatedly execute tasks based on a condition. Common use cases include:

  • API polling: Repeatedly call an API until it returns a specific status, such as SUCCESS, or until the response data meets your requirements.

  • Data readiness checks: Wait for an upstream file or data partition to be generated. The loop checks for the data in each iteration and exits once it is available, triggering downstream tasks.

  • Batch data processing: Process a list of items, such as table names or date partitions, provided by an upstream node like an assignment node. The do-while node iterates through the list, processing one item at a time until the list is exhausted.

  • Status synchronization: Repeatedly check the status of an external service. Once the service status changes to a desired state, such as "Available" or "Complete", the loop exits and triggers the subsequent data synchronization or processing workflow.

Usage notes

  • Version requirements: Available only in DataWorks Standard Edition and later.

  • Permissions: Your RAM account must be added to the target workspace and assigned the developer or workspace administrator role. For more information, see Add members to a workspace.

How it works

The do-while node acts as a loop container based on an "execute first, then check" mechanism.

image Start node → loop body (execute tasks) → image End node (evaluate condition)

  1. Start and execute: The loop starts at the image node. It executes all tasks in the loop body at least once.

  2. Evaluate the condition: After the loop body finishes, the image node runs. Write the condition logic in this node.

  3. Loop decision:

    • If the image node outputs the string True, a new iteration starts from step 1.

    • If the image node outputs the string False, the loop ends and the do-while node succeeds.

Key feature: The loop body always executes at least once, regardless of the initial condition.

Node components

Double-click a do-while node to open its internal canvas, which consists of three main parts:

  • image Start: Marks the entry point of the loop. You cannot edit or delete this node.

  • Loop body: A customizable canvas for your tasks. You can click Create Internal Node to add various types of internal nodes, such as Shell, SQL, and Python, to form the repeatable business process.

  • image End: Evaluates the loop condition. This node is an assignment node in which you can write ODPS SQL, Shell, or Python code. It outputs the string True or False to determine whether the loop continues.

Built-in variables

In the loop body and End node of a do-while node, you can use the following built-in variables to obtain the current loop state and input data:

Built-in variable

Description

System built-in variables

${dag.loopTimes}

The current iteration number, starting from 1.

${dag.offset}

The offset of the current iteration, starting from 0. Equivalent to ${dag.loopTimes} - 1.

Used with an assignment node
(Assuming the input parameter input is bound to the output of an upstream assignment node)

${dag.input}

Retrieves the complete result set passed from upstream (typically a two-dimensional array).

${dag.input.length}

Retrieves the number of rows in the result set (the total number of elements).

${dag.input[${dag.offset}]}

Retrieves the row of data being processed in the current iteration. The return format is a comma-separated string (for example, "user_A,beijing").

${dag.input[i][j]}

(For two-dimensional arrays only) Retrieves the exact value at row i, column j (indexes start from 0).

Example: ${dag.input[0][1]} extracts the value "beijing" from the first row and second column.

Notes

  • Edition requirement: This feature is available only in DataWorks Standard Edition and higher.

  • Loop limit: The default maximum number of iterations is 128. The configurable maximum is 1024. Exceeding this limit causes the task to fail.

  • Execution mechanism: The do-while node uses a serial execution mode and does not support concurrent processing. Each iteration must complete before the next one starts.

  • Debugging limitation: You cannot run a test directly on the Data Studio page. You must first deploy the workflow, and then use the backfill data feature in Operation Center to verify the task.

  • Dependency propagation: If the do-while node depends on an upstream assignment node, make sure to start the backfill from the upstream node to ensure the data pipeline is complete.

  • Flow control: If you use a Branch node inside the loop body, make sure all branch paths eventually converge at a Merge node to avoid breaking the workflow.

Procedure: Create a simple loop task

This example walks you through creating a task that loops 5 times and prints the current iteration number in each loop.

Step 1: Configure the loop body (add a Shell node)

  1. Create a do-while node in your workflow.

  2. Double-click the node to enter its internal canvas. In the loop body area, click Create Internal Node, select Shell, and name it print_loop_times.

  3. Right-click the node and select Open Node.

  4. Enter the following command in the code editor:

    # Use the built-in variable ${dag.loopTimes} to get the current loop count
    echo "This is loop number: ${dag.loopTimes}"
  5. Click the save icon to save the node.

Step 2: Define the exit condition

  1. Return to the internal canvas of the do-while node, right-click the image End node, and select Open Node.

  2. Switch the language to Python.

  3. Enter the following Python code:

    # When the loop count is less than 5, output True and continue looping
    # When the loop count reaches 5, output False and exit the loop
    if ${dag.loopTimes} < 5:
        print True
    else:
        print False
  4. Save the End node.

Step 3: Deploy, run, and verify

  1. Return to the main workflow canvas and click Deploy on the toolbar to deploy the entire workflow.

  2. Go to Operation Center and locate the do-while node.

  3. Right-click the node and choose Backfill Data > Current Node to start a test run.

  4. After the instance runs successfully, right-click the node again and select View Internal Nodes.

  5. View the five loop iterations. Expand any iteration (for example, the 5th), right-click the print_loop_times instance inside, and select View Runtime Log. You should see the following output:

    This is loop number: 5

Advanced usage: Process a data list with an assignment node

This is a common and practical scenario: use a do-while node to iterate through and process a dataset output by an upstream assignment node. In this example, an upstream ODPS SQL assignment node queries two rows of user information, and the do-while node processes each row.

Step 1: Configure the upstream assignment node

  1. Create an assignment node (for example, named assign_sql_data) and set it as the upstream of the do-while node.

  2. Use ODPS SQL to query data in the node:

    SELECT 'user_A', 'beijing'
    UNION ALL
    SELECT 'user_B', 'shanghai';
  3. Save the node. Its outputs parameter will output a two-dimensional array containing two rows of data.

Step 2: Configure the do-while node to consume data

  1. In the Schedule Settings panel on the right side of the do-while node, find Input Parameters of This Node.

  2. Click Add and configure the following parameter:

    • Parameter Name: input (customizable)

    • Value Source: Select assign_sql_data.outputs

  3. In the loop body of the do-while node, create a Shell node and enter the following script to process the current row:

    # Get the data row being processed in the current loop iteration
    echo "Processing data row: ${dag.input[${dag.offset}]}"
  4. Open the image End node and use Python to define the exit condition:

    # If the number of times executed is less than the total number of data rows, continue looping
    if ${dag.loopTimes} < ${dag.input.length}:
        print True
    else:
        print False
  5. Deploy the workflow, and in Operation Center, start the backfill from the assign_sql_data node to ensure data flows correctly into the loop.

After a successful run, you can see in the logs that the two iterations processed "user_A,beijing" and "user_B,shanghai" respectively.

Appendix: Comparison between do-while and for-each nodes

Feature

Do-while node

For-each node

Core logic

Condition-driven: Repeats execution until a condition is no longer met.

Data-driven: Executes once for each element in the input list.

Number of iterations

Indeterminate; depends on when the condition is met.

Determinate; equals the number of elements in the input list.

Execution guarantee

Executes at least once (even if the initial condition is already not met).

If the input is empty, no execution occurs.

Use cases

Polling, waiting, status checks, and batch processing until the source is empty.

Batch processing of a known list (such as synchronizing a set of tables or processing multiple partitions).

Control method

Controlled by the End node outputting True / False to determine whether the loop continues.

Automatically ends after all elements have been traversed.