Do-while node
DataWorks provides the do-while node, which implements "execute first, then check" loop logic. This node is ideal when you need to repeat a series of tasks until a dynamic condition is met, such as polling to check for a file's existence or processing data in batches until the source is empty. You can orchestrate a task flow within the loop body and use a dedicated End node to control the loop's exit.
Use cases
In data development, the do-while node simplifies the design of workflows that need to repeatedly execute tasks based on a condition. Common use cases include:
API polling: Repeatedly call an API until it returns a specific status, such as
SUCCESS, or until the response data meets your requirements.Data readiness checks: Wait for an upstream file or data partition to be generated. The loop checks for the data in each iteration and exits once it is available, triggering downstream tasks.
Batch data processing: Process a list of items, such as table names or date partitions, provided by an upstream node like an assignment node. The do-while node iterates through the list, processing one item at a time until the list is exhausted.
Status synchronization: Repeatedly check the status of an external service. Once the service status changes to a desired state, such as "Available" or "Complete", the loop exits and triggers the subsequent data synchronization or processing workflow.
Usage notes
Version requirements: Available only in DataWorks Standard Edition and later.
Permissions: Your RAM account must be added to the target workspace and assigned the developer or workspace administrator role. For more information, see Add members to a workspace.
How it works
The do-while node acts as a loop container based on an "execute first, then check" mechanism.
Start node → loop body (execute tasks) →
End node (evaluate condition)
Start and execute: The loop starts at the
node. It executes all tasks in the loop body at least once.Evaluate the condition: After the loop body finishes, the
node runs. Write the condition logic in this node.Loop decision:
If the
node outputs the string True, a new iteration starts from step 1.If the
node outputs the string False, the loop ends and the do-while node succeeds.
Key feature: The loop body always executes at least once, regardless of the initial condition.
Node components
Double-click a do-while node to open its internal canvas, which consists of three main parts:
Start: Marks the entry point of the loop. You cannot edit or delete this node.Loop body: A customizable canvas for your tasks. You can click Create Internal Node to add various types of internal nodes, such as Shell, SQL, and Python, to form the repeatable business process.
End: Evaluates the loop condition. This node is an assignment node in which you can write ODPS SQL, Shell, or Python code. It outputs the string TrueorFalseto determine whether the loop continues.
Built-in variables
In the loop body and End node of a do-while node, you can use the following built-in variables to obtain the current loop state and input data:
Built-in variable | Description |
System built-in variables | |
| The current iteration number, starting from 1. |
| The offset of the current iteration, starting from 0. Equivalent to |
Used with an assignment node | |
| Retrieves the complete result set passed from upstream (typically a two-dimensional array). |
| Retrieves the number of rows in the result set (the total number of elements). |
| Retrieves the row of data being processed in the current iteration. The return format is a comma-separated string (for example, |
| (For two-dimensional arrays only) Retrieves the exact value at row |
Example:${dag.input[0][1]}extracts the value"beijing"from the first row and second column.
Notes
Edition requirement: This feature is available only in DataWorks Standard Edition and higher.
Loop limit: The default maximum number of iterations is 128. The configurable maximum is 1024. Exceeding this limit causes the task to fail.
Execution mechanism: The do-while node uses a serial execution mode and does not support concurrent processing. Each iteration must complete before the next one starts.
Debugging limitation: You cannot run a test directly on the Data Studio page. You must first deploy the workflow, and then use the backfill data feature in Operation Center to verify the task.
Dependency propagation: If the do-while node depends on an upstream assignment node, make sure to start the backfill from the upstream node to ensure the data pipeline is complete.
Flow control: If you use a Branch node inside the loop body, make sure all branch paths eventually converge at a Merge node to avoid breaking the workflow.
Procedure: Create a simple loop task
This example walks you through creating a task that loops 5 times and prints the current iteration number in each loop.
Step 1: Configure the loop body (add a Shell node)
Create a
do-whilenode in your workflow.Double-click the node to enter its internal canvas. In the loop body area, click Create Internal Node, select Shell, and name it
print_loop_times.Right-click the node and select Open Node.
Enter the following command in the code editor:
# Use the built-in variable ${dag.loopTimes} to get the current loop count echo "This is loop number: ${dag.loopTimes}"Click the save icon to save the node.
Step 2: Define the exit condition
Return to the internal canvas of the do-while node, right-click the
End node, and select Open Node.Switch the language to Python.
Enter the following Python code:
# When the loop count is less than 5, output True and continue looping # When the loop count reaches 5, output False and exit the loop if ${dag.loopTimes} < 5: print True else: print FalseSave the End node.
Step 3: Deploy, run, and verify
Return to the main workflow canvas and click Deploy on the toolbar to deploy the entire workflow.
Go to Operation Center and locate the
do-whilenode.Right-click the node and choose Backfill Data > Current Node to start a test run.
After the instance runs successfully, right-click the node again and select View Internal Nodes.
View the five loop iterations. Expand any iteration (for example, the 5th), right-click the
print_loop_timesinstance inside, and select View Runtime Log. You should see the following output:This is loop number: 5
Advanced usage: Process a data list with an assignment node
This is a common and practical scenario: use a do-while node to iterate through and process a dataset output by an upstream assignment node. In this example, an upstream ODPS SQL assignment node queries two rows of user information, and the do-while node processes each row.
Step 1: Configure the upstream assignment node
Create an assignment node (for example, named
assign_sql_data) and set it as the upstream of thedo-whilenode.Use ODPS SQL to query data in the node:
SELECT 'user_A', 'beijing' UNION ALL SELECT 'user_B', 'shanghai';Save the node. Its
outputsparameter will output a two-dimensional array containing two rows of data.
Step 2: Configure the do-while node to consume data
In the Schedule Settings panel on the right side of the
do-whilenode, find Input Parameters of This Node.Click Add and configure the following parameter:
Parameter Name:
input(customizable)Value Source: Select
assign_sql_data.outputs
In the loop body of the
do-whilenode, create a Shell node and enter the following script to process the current row:# Get the data row being processed in the current loop iteration echo "Processing data row: ${dag.input[${dag.offset}]}"Open the
End node and use Python to define the exit condition:# If the number of times executed is less than the total number of data rows, continue looping if ${dag.loopTimes} < ${dag.input.length}: print True else: print FalseDeploy the workflow, and in Operation Center, start the backfill from the
assign_sql_datanode to ensure data flows correctly into the loop.
After a successful run, you can see in the logs that the two iterations processed "user_A,beijing" and "user_B,shanghai" respectively.
Appendix: Comparison between do-while and for-each nodes
Feature | Do-while node | For-each node |
Core logic | Condition-driven: Repeats execution until a condition is no longer met. | Data-driven: Executes once for each element in the input list. |
Number of iterations | Indeterminate; depends on when the condition is met. | Determinate; equals the number of elements in the input list. |
Execution guarantee | Executes at least once (even if the initial condition is already not met). | If the input is empty, no execution occurs. |
Use cases | Polling, waiting, status checks, and batch processing until the source is empty. | Batch processing of a known list (such as synchronizing a set of tables or processing multiple partitions). |
Control method | Controlled by the End node outputting | Automatically ends after all elements have been traversed. |