Configure a scheduling policy

Updated at:

A scheduling policy defines how task instances are generated, executed, retried on failure, and allocated compute resources. You can configure a scheduling policy to control instance generation timing, execution logic, failure recovery, and resource allocation for stable data development and operations.

Key concepts

A DataWorks scheduling policy controls the automated behavior of a scheduled task across four aspects:

  1. Instance generation and execution
    Control when an instance is generated (the next day or immediately) and how it runs (normal, paused, or dry run).

  2. Fault tolerance and retries
    Configure how to handle exceptions such as runtime timeouts, automatic retries on failure, and permissions for manual reruns.

  3. Resources and environment
    Specify the compute resources, resource group, and custom image for a task to achieve resource isolation and a consistent environment.

  4. Concurrency control
    Set the maximum number of concurrent instances to prevent resource overload.

image

Configure a scheduling policy

Instance generation and execution

Configure when and how task instances are created and executed.

Instance generation mode

Specify when scheduled instances are generated after a task is deployed. This setting determines whether changes take effect on the current day or the next day.

Note

The scope of scheduling policy configurations varies based on the node's organizational structure:

  • Workflow node: The Instance generation method is centrally managed by the workflow, and the node itself cannot be configured independently. Other scheduling parameters can be set independently for the node.

  • Standalone nodes: All scheduling policy parameters can be configured independently.

Instance generation mode

Description

T +1 generated next day (recommended)

Changes take effect the next day and do not affect instances already generated for the current day.

  • New task: After being deployed, the task is automatically scheduled to run starting the next day. To run the task on the current day, you must manually create an instance using the backfill data feature.

  • Existing task modification: Changes take effect the next day. Instances already generated for the current day are not affected.

Note

If you modify the scheduling time, select this option, and deploy the task, the scheduled time for instances already generated for the current day (T) and the previous day (T-1), including completed and unrun instances, will be updated to the new time. Future instances that have not yet been generated will be created based on the new time.

Instant generation after publishing

Changes take effect immediately after deployment, and instances for the current day are regenerated. If the scheduled time is less than 10 minutes after the deployment time, the newly generated instance will automatically perform a dry run (marked as successful without actual execution).

  • New task: Whether the task actually runs on the deployment day depends on the scheduled time. If the interval between the scheduled time and the deployment time is less than 10 minutes, the instance will perform a dry run.

  • Existing task with schedule modification: The system regenerates future instances based on the new configuration. Historical instances that have already been generated are not affected.

Important
  • If you select Instant generation after publishing, to ensure the instance runs normally instead of performing a dry run, the scheduled time must be at least 10 minutes later than the deployment time.

  • Regardless of the selected mode, if a task is deployed between 22:00~24:00, the changes will take effect on the third day. We recommend that you avoid deploying tasks during this period.

Scheduling type

Specify how an instance behaves when the scheduled time is reached.

Scheduling type

Description

Use case

Normal scheduling

The instance executes the code normally and triggers downstream tasks upon success.

Suitable for all regular tasks that run on a periodic basis.

Pause scheduling

When the scheduled time is reached, the instance does not run and is marked as failed, which blocks downstream dependencies. In Operation Center, the node displays a freeze icon.

Suitable for temporarily interrupting or undeploying a business process, such as urgently cutting off a data pipeline.

Empty Run Scheduling

When the scheduled time is reached, the instance is directly marked as successful (with a duration of 0 seconds). The code is not actually executed and no compute resources are consumed, but downstream tasks are triggered normally.

Suitable for nodes that temporarily do not need to run but whose downstream tasks must still be triggered.

Delayed execution time (workflow nodes only)

Specify the delay before a node starts running after the workflow begins.

Example: A workflow has a scheduled time of 09:00, and one of its nodes is configured with a 5-minute delayed execution. The node will start running at approximately 09:05.

Fault tolerance and retries

Configure strategies for handling task failures and stalls.

Timeout

Set the maximum allowed runtime for a task. If the task exceeds this threshold, it is forcibly terminated and marked as failed to prevent a stalled task from blocking the workflow.

  • System Default: The default value is 3 to 7 days, dynamically adjusted by the system based on the current load.

  • Custom: You can set a timeout from 1 minute to 168 hours (7 days).

Note

The timeout setting applies to scheduled instances, backfill instances, and test instances.

Rerun property

Specify whether a task can be manually rerun after execution. Use this to enable fault recovery or prevent data quality issues from accidental reruns.

Rerun property

Description and use case

You can run again after success or failure.

Applicable to idempotent tasks that produce the same result regardless of how many times they run.

After successful operation, you cannot run again. After failed operation, you can run again.

(Recommended) Applicable to tasks where rerunning after success may cause side effects such as duplicate data insertion. This prevents data pollution from accidentally rerunning successful tasks.

Do not run again after success or failure.

Applicable to strictly non-idempotent tasks such as one-time data synchronization scenarios. When this option is selected, the Auto Rerun upon Failure feature cannot be enabled.

Note

To ensure data recoverability, we recommend that you make your task logic idempotent. For example, in ODPS SQL tasks, use INSERT OVERWRITE (overwrite) instead of INSERT INTO (append) to ensure consistent results across multiple reruns.

Auto rerun upon failure

When enabled, the system automatically retries a task that fails due to transient issues such as network jitter, without manual intervention.

Parameter

Description

Number of Reruns

The number of automatic retries after a failure. Valid values: 1 to 10.

Rerun interval

The interval between retries. Valid values: 1 to 30 minutes.

Important

Failures caused by a runtime timeout do not trigger the auto rerun mechanism.

Runtime environment and resources

Configure the compute resources, runtime environment, and external data for task execution.

Compute resources and quotas

For information about compute resource management, see Compute Resources Overview.

Compute resources are compute engine instances that perform data processing tasks, such as MaxCompute projects, Hologres instances, AnalyticDB, and ClickHouse.

For MaxCompute compute resources, you can use the Computing Quota mechanism in the MaxCompute console to isolate and allocate compute resources. For example, you can create an independent quota group for critical workloads and set CPU, memory, and concurrency limits to prevent resource contention from affecting core tasks.

Resource group

For information about resource group management, see Resource Group Overview.

All scheduled tasks in DataWorks require a resource group for scheduling and execution. Tasks fall into two types based on whether they consume DataWorks compute units (CUs):

Task type

Core function

CU consumption

Task list

Scheduling tasks

Triggers and monitors compute jobs in external compute engines such as MaxCompute and Hologres.

No

Scheduling task configuration list

Compute tasks

Executes task code directly within the DataWorks resource group.

Yes

CU configuration list for compute tasks

  1. In the Resource Properties section of schedule settings, select a resource group from the Scheduling Resource Group drop-down list. We recommend that you use a serverless resource group, which offers greater versatility, elastic scaling, and flexible billing.

    Important

    Make sure that the selected resource group has network connectivity to the data sources and compute resources that the task needs to access. Otherwise, the task will fail. For information about network connectivity configuration, see Network connectivity configuration.

  2. If the task is a compute task and you selected a serverless resource group in the previous step, you must configure the scheduling CU. Based on the computational complexity and performance requirements of the task, set the number of CUs required at runtime (for example, 0.25 CU, 1 CU, or 4 CU).

Image

For information about custom image management, see Custom image.

Select a custom runtime image for specific node types such as Python or Shell. Custom images let you pre-install third-party libraries and tools, providing a consistent execution environment for tasks.

Dataset

For information about dataset management, see Manage datasets

Mount external storage, such as OSS or File Storage (NAS), for specific node types such as Shell, so that the data can be read as local files in the code. Each node supports up to 5 datasets.

You can configure the Mount Path, Advanced Settings (such as Read Method), and Read Only permissions. For information about how to configure datasets, see Manage datasets.

Concurrency control

Control concurrent execution of multiple instances of the same task to prevent resource contention or business logic conflicts.

Maximum concurrent instances

Set the maximum number of instances of the same task that can run concurrently.

  • Allowed: The default option. No limit is imposed on the number of concurrent instances for the task.

  • Limit: When enabled, you can set the maximum concurrency (valid values: 1 to 10000). When the number of running instances reaches the limit, newly generated instances enter a queue and run sequentially after existing instances complete.

Billing

  • Scheduling tasks: Only trigger external compute engines and do not consume DataWorks compute units (CUs).

  • Compute tasks: If executed in DataWorks resource groups (including default and serverless resource groups), CUs are consumed and fees are incurred.

    • When using a serverless resource group, you must configure the scheduling CU specification based on task requirements. Fees are based on CU usage and duration.

  • Underlying compute resource fees: The scheduling policy covers only DataWorks-level resource scheduling. The underlying compute resources used during task execution, such as MaxCompute and Hologres, are billed separately. For more information, see the billing documentation of the corresponding product.

FAQ

Does an exhausted OpenAPI quota affect scheduled task execution?

No. Scheduled instances are generated in advance by the system and do not rely on OpenAPI calls. If your OpenAPI quota is exhausted, scheduled tasks that are deployed to the production environment continue to generate and run instances as planned. Only O&M operations initiated through API calls are affected.