Configure a scheduling policy
A scheduling policy defines how task instances are generated, executed, retried on failure, and allocated compute resources. You can configure a scheduling policy to control instance generation timing, execution logic, failure recovery, and resource allocation for stable data development and operations.
Key concepts
A DataWorks scheduling policy controls the automated behavior of a scheduled task across four aspects:
-
Instance generation and execution
Control when an instance is generated (the next day or immediately) and how it runs (normal, paused, or dry run). -
Fault tolerance and retries
Configure how to handle exceptions such as runtime timeouts, automatic retries on failure, and permissions for manual reruns. -
Resources and environment
Specify the compute resources, resource group, and custom image for a task to achieve resource isolation and a consistent environment. -
Concurrency control
Set the maximum number of concurrent instances to prevent resource overload.
Configure a scheduling policy
Instance generation and execution
Configure when and how task instances are created and executed.
Instance generation mode
Specify when scheduled instances are generated after a task is deployed. This setting determines whether changes take effect on the current day or the next day.
The scope of scheduling policy configurations varies based on the node's organizational structure:
-
Workflow node: The Instance generation method is centrally managed by the workflow, and the node itself cannot be configured independently. Other scheduling parameters can be set independently for the node.
-
Standalone nodes: All scheduling policy parameters can be configured independently.
|
Instance generation mode |
Description |
|
T +1 generated next day (recommended) |
Changes take effect the next day and do not affect instances already generated for the current day.
Note
If you modify the scheduling time, select this option, and deploy the task, the scheduled time for instances already generated for the current day (T) and the previous day (T-1), including completed and unrun instances, will be updated to the new time. Future instances that have not yet been generated will be created based on the new time. |
|
Instant generation after publishing |
Changes take effect immediately after deployment, and instances for the current day are regenerated. If the scheduled time is less than 10 minutes after the deployment time, the newly generated instance will automatically perform a dry run (marked as successful without actual execution).
|
-
If you select Instant generation after publishing, to ensure the instance runs normally instead of performing a dry run, the scheduled time must be at least 10 minutes later than the deployment time.
-
Regardless of the selected mode, if a task is deployed between
22:00~24:00, the changes will take effect on the third day. We recommend that you avoid deploying tasks during this period.
Scheduling type
Specify how an instance behaves when the scheduled time is reached.
|
Scheduling type |
Description |
Use case |
|
Normal scheduling |
The instance executes the code normally and triggers downstream tasks upon success. |
Suitable for all regular tasks that run on a periodic basis. |
|
Pause scheduling |
When the scheduled time is reached, the instance does not run and is marked as failed, which blocks downstream dependencies. In Operation Center, the node displays a freeze icon. |
Suitable for temporarily interrupting or undeploying a business process, such as urgently cutting off a data pipeline. |
|
Empty Run Scheduling |
When the scheduled time is reached, the instance is directly marked as successful (with a duration of |
Suitable for nodes that temporarily do not need to run but whose downstream tasks must still be triggered. |
Delayed execution time (workflow nodes only)
Specify the delay before a node starts running after the workflow begins.
Example: A workflow has a scheduled time of 09:00, and one of its nodes is configured with a 5-minute delayed execution. The node will start running at approximately 09:05.
Fault tolerance and retries
Configure strategies for handling task failures and stalls.
Timeout
Set the maximum allowed runtime for a task. If the task exceeds this threshold, it is forcibly terminated and marked as failed to prevent a stalled task from blocking the workflow.
-
System Default: The default value is 3 to 7 days, dynamically adjusted by the system based on the current load.
-
Custom: You can set a timeout from 1 minute to 168 hours (7 days).
The timeout setting applies to scheduled instances, backfill instances, and test instances.
Rerun property
Specify whether a task can be manually rerun after execution. Use this to enable fault recovery or prevent data quality issues from accidental reruns.
|
Rerun property |
Description and use case |
|
You can run again after success or failure. |
Applicable to idempotent tasks that produce the same result regardless of how many times they run. |
|
After successful operation, you cannot run again. After failed operation, you can run again. |
(Recommended) Applicable to tasks where rerunning after success may cause side effects such as duplicate data insertion. This prevents data pollution from accidentally rerunning successful tasks. |
|
Do not run again after success or failure. |
Applicable to strictly non-idempotent tasks such as one-time data synchronization scenarios. When this option is selected, the Auto Rerun upon Failure feature cannot be enabled. |
To ensure data recoverability, we recommend that you make your task logic idempotent. For example, in ODPS SQL tasks, use INSERT OVERWRITE (overwrite) instead of INSERT INTO (append) to ensure consistent results across multiple reruns.
Auto rerun upon failure
When enabled, the system automatically retries a task that fails due to transient issues such as network jitter, without manual intervention.
|
Parameter |
Description |
|
Number of Reruns |
The number of automatic retries after a failure. Valid values: 1 to 10. |
|
Rerun interval |
The interval between retries. Valid values: 1 to 30 minutes. |
Failures caused by a runtime timeout do not trigger the auto rerun mechanism.
Runtime environment and resources
Configure the compute resources, runtime environment, and external data for task execution.
Compute resources and quotas
For information about compute resource management, see Compute Resources Overview.
Compute resources are compute engine instances that perform data processing tasks, such as MaxCompute projects, Hologres instances, AnalyticDB, and ClickHouse.
For MaxCompute compute resources, you can use the Computing Quota mechanism in the MaxCompute console to isolate and allocate compute resources. For example, you can create an independent quota group for critical workloads and set CPU, memory, and concurrency limits to prevent resource contention from affecting core tasks.
Resource group
For information about resource group management, see Resource Group Overview.
All scheduled tasks in DataWorks require a resource group for scheduling and execution. Tasks fall into two types based on whether they consume DataWorks compute units (CUs):
|
Task type |
Core function |
CU consumption |
Task list |
|
Scheduling tasks |
Triggers and monitors compute jobs in external compute engines such as MaxCompute and Hologres. |
No |
|
|
Compute tasks |
Executes task code directly within the DataWorks resource group. |
Yes |
-
In the Resource Properties section of schedule settings, select a resource group from the Scheduling Resource Group drop-down list. We recommend that you use a serverless resource group, which offers greater versatility, elastic scaling, and flexible billing.
ImportantMake sure that the selected resource group has network connectivity to the data sources and compute resources that the task needs to access. Otherwise, the task will fail. For information about network connectivity configuration, see Network connectivity configuration.
-
If the task is a compute task and you selected a serverless resource group in the previous step, you must configure the scheduling CU. Based on the computational complexity and performance requirements of the task, set the number of CUs required at runtime (for example, 0.25 CU, 1 CU, or 4 CU).
Image
For information about custom image management, see Custom image.
Select a custom runtime image for specific node types such as Python or Shell. Custom images let you pre-install third-party libraries and tools, providing a consistent execution environment for tasks.
Dataset
For information about dataset management, see Manage datasets
Mount external storage, such as OSS or File Storage (NAS), for specific node types such as Shell, so that the data can be read as local files in the code. Each node supports up to 5 datasets.
You can configure the Mount Path, Advanced Settings (such as Read Method), and Read Only permissions. For information about how to configure datasets, see Manage datasets.
Concurrency control
Control concurrent execution of multiple instances of the same task to prevent resource contention or business logic conflicts.
Maximum concurrent instances
Set the maximum number of instances of the same task that can run concurrently.
-
Allowed: The default option. No limit is imposed on the number of concurrent instances for the task.
-
Limit: When enabled, you can set the maximum concurrency (valid values: 1 to 10000). When the number of running instances reaches the limit, newly generated instances enter a queue and run sequentially after existing instances complete.
Billing
-
Scheduling tasks: Only trigger external compute engines and do not consume DataWorks compute units (CUs).
-
Compute tasks: If executed in DataWorks resource groups (including default and serverless resource groups), CUs are consumed and fees are incurred.
-
When using a serverless resource group, you must configure the scheduling CU specification based on task requirements. Fees are based on CU usage and duration.
-
-
Underlying compute resource fees: The scheduling policy covers only DataWorks-level resource scheduling. The underlying compute resources used during task execution, such as MaxCompute and Hologres, are billed separately. For more information, see the billing documentation of the corresponding product.
FAQ
Does an exhausted OpenAPI quota affect scheduled task execution?
No. Scheduled instances are generated in advance by the system and do not rely on OpenAPI calls. If your OpenAPI quota is exhausted, scheduled tasks that are deployed to the production environment continue to generate and run instances as planned. Only O&M operations initiated through API calls are affected.