A scheduling cycle defines how often a node automatically runs in the production environment. Based on this cycle, the DataWorks scheduling system generates recurring instances and runs them according to node dependencies and the scheduled time of each instance.
Key concepts
-
Recurring instance: For each data timestamp, the scheduling system generates a runtime entity for an auto triggered task based on its scheduling configuration, such as running at 00:00 every day. This entity is a recurring instance. The node's runtime, status, and logs are associated with this instance.
-
Cross-cycle dependency: In DataWorks, nodes with different scheduling cycles can depend on each other. For example, a daily downstream node can depend on an hourly upstream node. The dependency between nodes is a dependency between their recurring instances. For more information, see Cross-cycle dependencies.
-
Dry-run: For nodes that do not run daily, such as weekly, monthly, or yearly nodes, the scheduling system generates a dry-run instance on non-scheduled days. When this instance reaches its scheduled time, its status immediately changes to "Success", but it does not run the code logic within the node. The main purpose of a dry-run is to resolve dependencies and ensure that downstream daily nodes can be triggered correctly.
-
The instance status is "Success", the running time is 0 seconds, and no logs are generated.
-
It does not use scheduling or computing resources.
-
It does not block the execution of downstream nodes. Even if an upstream node is a dry-run instance, downstream nodes run as scheduled after their running conditions are met.
-
Basic principles and scenarios
Conditions for scheduled execution
A recurring instance runs only after both of the following conditions are met. The conditions can be met in any order.
-
All its upstream instances have run successfully. This includes instances that completed as a dry-run.
-
The instance has reached its scheduled time.
Therefore, the configured scheduled time is only the expected scheduled time. The actual running time of a node is affected by factors such as the completion time of its upstream nodes, available computing resources, and the node's specific running conditions.
The period around 00:00 each day is the peak scheduling time for DataWorks, when a large number of tasks are triggered simultaneously. If your task is scheduled at 00:00, there is typically a delay of several minutes between the scheduled time and the actual start of execution. This is expected behavior.
If your task is sensitive to its start time, we recommend that you set the scheduled time to a time after 00:00 to avoid the peak period (for example, 00:10 or 01:00). If your business requires the task to start at 00:00 (such as the first-layer task in a T+1 data warehouse processing pipeline), account for this delay in your downstream timeliness assessment.
Workflow scheduling scenarios
In a workflow that consists of nodes A, B, and C (A→B→C), the scheduled time configuration affects the execution of the entire workflow:
|
Scenario 1: Unified start time |
|
|
|
The entire workflow must start after 03:00. To achieve this, set the scheduled time of the start node A to |
|
Scenario 2: Different start times for each node |
|
|
|
Node A is scheduled to run at 03:00, node B must not run until after 05:00, and node C must not run until after 06:00. You need to set the scheduled times of nodes A, B, and C to |
|
Scenario 3: Specific start times for some nodes |
|
|
|
Node A is scheduled to run at 03:00, node B must not run until after 05:00, and node C has no specific time requirement. Set the scheduled time of node A to |
Impact of updating the scheduled time
After you modify the scheduled time of a node in Scheduling Settings and redeploy the node, the impact depends on the selected instance generation mode:
-
Mode A: Next day (T+1, default): After you select this option and deploy the node, the scheduled times of instances that have already been generated for the last two days (T and T-1), including completed and pending instances, are updated to the new configured time. Future instances that have not yet been generated will be generated with the new time.
-
Mode B: Immediately after deployment: Selecting this option immediately generates new instances based on the new configuration. The scheduled times of historical instances remain unchanged.
Scheduling type configuration
DataWorks supports six scheduling cycles: minute, hour, day, week, month, and year. On the node editing page in DataStudio, click Scheduling Settings on the right side, and configure the settings in the Scheduling time section.
The way you configure the scheduled time depends on how the node is organized:
-
Workflow node: If a node is organized within a workflow, the scheduled time is configured at the workflow level. Individual nodes within the workflow cannot have their own scheduled time configurations. To modify the scheduled time, go to the scheduling configuration of the workflow.
-
Standalone node: If a node does not belong to any workflow, the scheduled time is configured independently on the node itself.
Minute-level scheduling
The minimum interval for minute-level scheduling is 1 minute.
Configuration example
Goal: The node is scheduled every 30 minutes between 00:00 and 23:59 every day.
The cron expression is automatically generated based on the selected time and cannot be manually modified.
Set the effective date to Permanent. The automatically generated cron expression is 00 */30 00-23 * * ?.
Scheduling details
The preceding configuration generates 48 recurring instances per day, with scheduled times of 00:00, 00:30, 01:00, ... , 23:30. The data timestamp (bizdate) of each instance is the current day.
For more dependency scenarios of minute-level scheduling, see Cross-cycle dependencies for minute-level scheduling.
Hourly scheduling
Notes
-
The time range follows the inclusive-inclusive principle. For example, if you configure a node to run every 1 hour between
00:00and03:00, the scheduling system generates 4 instances with scheduled times of00:00,01:00,02:00, and03:00. -
You can set a time range and interval, or directly specify multiple specific run times.
Configuration example
Goal: The node is scheduled every 6 hours between 00:00 and 23:59 every day.
The cron expression is automatically generated based on the selected time and cannot be manually modified.
Set Scheduling Cycle to Hour, select the Hour Range tab, set Effective Date to Permanent. The automatically generated cron expression is 00 00 00-23/6 * * ?.
Scheduling details
The scheduling system generates 4 instances per day, with scheduled times of 00:00, 06:00, 12:00, and 18:00.
For more dependency scenarios of hourly scheduling, see Cross-cycle dependencies for hourly scheduling.
Daily scheduling
A daily-scheduled node runs once per day at the specified scheduled time. When you create a node, the default scheduled time is randomly generated within the 00:00–00:30 range to prevent a large number of tasks from starting at 00:00 simultaneously. You can modify this as needed.
Configuration example
Goal: The node runs once per day at 13:00.
The cron expression is automatically generated based on the selected time and cannot be manually modified.
Set Effective Date to Permanent.
Scheduling details
The scheduling system generates one instance per day for this task, with a scheduled time of 13:00 on the current day.
For more dependency scenarios of daily scheduling, see Cross-cycle dependencies for daily scheduling.
Weekly scheduling
On non-scheduled days, a weekly-scheduled node triggers a dry-run to ensure that downstream dependencies can run normally. For more information, see Dry-run.
Configuration example
Goal: The task runs at the specified time every Monday and Friday. Instances generated on Monday and Friday are scheduled and run normally, while instances on other days are dry-run instances.
The cron expression is automatically generated based on the selected time and cannot be manually modified.
Set Scheduling time to 00:00, set Effective Date to Permanent. The automatically generated cron expression is 00 00 00 ? * 1,5.
Scheduling details
The scheduling system automatically generates and runs instances for the task.
When you use the backfill data feature, pay attention to the selected data timestamp. In DataWorks, data timestamp = scheduled date - 1. For example, to backfill a weekly-scheduled task that runs on Monday, select Sunday of the previous week as the data timestamp. If you select other dates, the backfill instances will be dry-run instances.
Monthly scheduling
On non-scheduled days, a monthly node triggers a dry-run to ensure that downstream dependencies can run properly. For more information, see Dry-run.
Monthly scheduling allows you to set Specified time to Last day of every month.
Configuration example
Goal: The node runs at a specified time on the last day of every month. The instance generated on the last day of each month runs as scheduled, and instances on other days are dry-runs.
The cron expression is automatically generated based on the selected time and cannot be manually modified.
Set Scheduling Cycle to Month, and set the specified time to 00:00 on the last day of every month.
Scheduling details
The scheduling system automatically generates and runs instances for the node.
When you use the backfill data feature, pay attention to the selected data timestamp. In DataWorks, data timestamp = scheduled date - 1. For example, to backfill a month-end node that runs on January 31, set the data timestamp to January 30. If you select a different date, the backfill instance will be a dry-run.
Yearly scheduling
On non-scheduled days, a yearly node triggers a dry-run to ensure that downstream dependencies can run properly. For more information, see Dry-run.
Configuration example
Goal: The node runs on the 1st and last day of January, April, July, and October every year. Instances generated on the specified dates run as scheduled, and instances on other dates are dry-runs.
In the form, set Scheduling time to 00:00, and set Effective Date to Permanent. The corresponding cron expression is 00 00 00 L,1 1,4,7,10 ?.
Scheduling details
The scheduling system automatically generates and runs instances for the node.
For more dependency scenarios, see Cross-cycle dependencies.