Scheduled workflows automate recurring data processing by generating task instances on a preset schedule (daily, monthly, and more). Each task runs only when its scheduled time arrives and all upstream dependencies are met, keeping complex data pipelines stable and orderly. Typical use cases include:
Automate recurring data processing: Synchronize, cleanse, or aggregate data at daily, hourly, or weekly intervals.
Build complex DAG dependency flows: Visually integrate nodes such as MaxCompute SQL, Hologres, EMR, and Python, define upstream and downstream dependencies, and enable automated scheduling.
Centrally manage and schedule multiple subtasks: Group logically related tasks into a single workflow to schedule, maintain, and monitor them as one unit.
Quick start
This feature is available in the new version of DataWorks Data Studio. To learn how to distinguish between the new and old versions, see Distinguish between the new and legacy Data Studio.
This section walks through a ready-to-run scheduled workflow. You will build a simple pipeline where a virtual node (the starting point) triggers a MaxCompute SQL node (for data processing). The workflow automatically calculates the total number of orders from the previous day and writes the results to a table each morning.
Step 1: Prepare the compute engine and data
In the target workspace, bind a MaxCompute compute engine.
In MaxCompute, create the following table to store the results.
-- Create a simple result table CREATE TABLE IF NOT EXISTS dw_order_count_test ( order_date STRING, total_count BIGINT ) PARTITIONED BY (ds STRING); -- Partitioned by ds to store daily aggregation results
Step 2: Create a scheduled workflow
Go to the Workspaces page in the DataWorks console. In the top navigation bar, select a desired region. Find the desired workspace and choose in the Actions column.
If the button is labeled Data Development, it opens the legacy Data Studio. Do not click it.
Click the
icon in the left navigation bar, and to the right of Project Directory, click to open the Create Workflow page.In the Create Workflow dialog box, set Scheduling Type to Periodic Scheduling, enter the required information (for example, set Name to
minimal_daily_demo), and complete the creation.
Step 3: Orchestrate the workflow: Drag nodes and connect dependencies
On the workflow canvas, drag a Zero-Load Node from the left component panel and name it
start_node.Zero-Load Node is used only to define the starting point of a business process and does not actually run.
Drag a MaxCompute SQL node and name it
count_orders.Click the circle at the bottom of
start_nodeand drag a line to the top ofcount_ordersto build a simple processing pipeline.
Step 4: Develop the node code
We recommend that you enable Data Agent to get intelligent code completion suggestions and improve development efficiency.
Double-click the
count_ordersnode to open the node code editor.Write the business logic code for the node (simulated statistical data is used here).
-- bizdate is a defined scheduling variable whose meaning needs to be specified in the scheduling configuration INSERT OVERWRITE TABLE dw_order_count_test PARTITION (ds='${bizdate}') SELECT '${bizdate}' as order_date, COUNT(*) as total_count FROM (SELECT 1 as id UNION ALL SELECT 2 as id) t; -- Simulated dataFor more information about node development, see Develop a MaxCompute SQL node.
Click the Save button at the top of the node editor to save the configuration.
Step 5: Configure schedule and parameters
Go back to the workflow. On the right side of the workflow canvas, click Scheduling Settings > Scheduling time tab:
Set Scheduling Frequency to Day.
Set Scheduling time to
00:05(that is, 00:05 every day).
On the right side of the
count_ordersnode editor, configure . Add a parameter with Parameter name set tobizdateand Parameter Value set to$[yyyymmdd-1](this represents the current date minus one day, that is, the previous day).
Step 6: Debug a single node and the entire workflow
count_ordersnode debugging:Configure debug parameters: Click Debug Configuration on the right side of the node editing page.
In Compute Resource, select the MaxCompute compute resource prepared in Step 1.
In Script Parameters, enter Value Used in This Run. The default value is the day before the current date.
Run the debug task: Click the Run button in the toolbar. The node runs with the debug parameters you configured in Debug Configuration.
After the run results meet expectations, click Sync to Scheduling in the upper-right corner to synchronize the run configuration to the schedule settings.
Workflow debugging:
Go back to the workflow canvas and click the
icon in the top toolbar.In the dialog box that appears, enter Value Used in This Run for the workflow (for example, if today is 20260120,
bizdateshould be replaced with20260119).
Step 7: Deploy to production
Go back to the workflow and click the
button in the top toolbar.In the deployment panel, the system performs dependency and configuration checks. After you confirm that everything is correct, click Start Release Production and set the deployment method to Full Publishing.
After a successful deployment, go to Operation and Maintenance Center to check whether the workflow appears in the scheduled task list.
You have completed the development of a simple scheduled workflow. This workflow runs automatically every day in the early morning.
Core design and configuration
Workflow orchestration uses a visual DAG canvas to organize tasks with control nodes (such as join and branch nodes) and interaction nodes (such as HTTP triggers), pass context through scheduling parameters, and define execution order and trigger conditions through scheduling dependencies.
Node/workflow orchestration
Simple process orchestration
Data development typically involves complex pipelines from multi-source integration to layered modeling (such as ODS and DWD layer construction). DataWorks uses visual orchestration to break down complex logic into functional sub-nodes and build standardized processing pipelines. This directed acyclic graph (DAG)-based model enables state-driven automated flow: when an upstream node succeeds, it immediately triggers downstream tasks, ensuring linear, stable, and orderly end-to-end processing. At this stage, the orchestration is the simplest form — a static, linear, and non-reversible DAG.
Complex orchestration: Flow control
Branch/join nodes and for-each/do-while nodes are available only in DataWorks Standard Edition and later.
Flow control nodes elevate data development from task integration to business orchestration. They go beyond the single linear dependency model of traditional DAGs by adding advanced logic through a set of precise control nodes.
Node name | Node description |
Virtual node | A virtual node does not perform actual computations. It centrally manages multiple subtasks and serves as the starting node of a workflow. For example, in a product order analysis workflow, a virtual node named Note For standalone node development, you need to use the workspace root node as the starting dependency node. |
Branch node | Routes to different downstream logic based on upstream results. For example, if the daily order total is 0, trigger an alert node and stop subsequent computations; if the total is greater than 0, continue to the report generation node. |
Join node | Merges the execution results of multiple branches to resolve downstream dependency issues. For example, a financial settlement task depends on two branches: "normal settlement" and "adjustment logic". The join node ensures that the final report archiving is triggered as long as either branch completes successfully. |
For-each node | Iterates over the result set from an assignment node and runs downstream operations on each element. For example, for 31 province names obtained by an assignment node, the for-each node runs a data cleansing task 31 times, processing one province's data partition each time. |
Do-while node | Repeats execution until a condition is met. For example, call an external API every 10 minutes to query the data synchronization status. If the returned value is "Processing", the loop continues. If the returned value is "Completed", the loop exits and subsequent processing starts. |
For more information, see Common nodes in Data Studio.
Complex orchestration: State awareness and external integration
Check nodes are available only in DataWorks Professional Edition and later. Other nodes are available only in DataWorks Enterprise Edition.
These nodes evaluate whether required physical resources or preceding tasks are ready, and integrate with third-party systems to bridge communication between data platforms and business systems.
Node name | Node description |
HTTP trigger | Receives HTTP requests from external systems to trigger DataWorks tasks. For example, after an upstream business system completes daily closing, it calls an HTTP API to trigger the DataWorks T+1 data processing pipeline. |
Check node | Monitors whether external resources (such as OSS files or MaxCompute partitions) are ready, and triggers downstream tasks once they become available. For example, a check node waits for the day's log file to appear in OSS, then starts the log parsing task. |
Dependency check node | Uses active polling, logic combinations, and cross-cycle or cross-workspace dependency readiness checks to trigger downstream tasks after all conditions are met. For example, a daily scheduled node waits for all 24 hourly tasks from the previous day to complete. |
For more information, see Common nodes in Data Studio.
Encapsulate sub-workflows for logic reuse
You can encapsulate a stable, reusable sub-task pipeline into a sub-workflow through a SUB_PROCESS node, which other workflows can then reference. For example, e-commerce, advertising, and IoT business lines can share a standardized pipeline for data cleansing, summary statistics, and Data Quality checks.
Procedure:
In the workflow
child_workflowthat you plan to reference, set the workflow General property to Can be cited. This converts the workflow into a sub-workflow.In the main workflow, drag a SUB_PROCESS node and set the referenced workflow to
child_workflow.
Sub-workflows have the following constraints:
Internality: The workflow and all its internal nodes cannot have any dependencies on external tasks.
Isolation: The workflow cannot be directly set as a dependency by any external task.
Passive triggering: After deployment, the workflow does not automatically generate scheduled instances. It runs only when called by a
SUB_PROCESSnode in another workflow.For more information, see Create a sub-workflow.
Workflow splitting and modular design recommendations
To keep workflows maintainable and performant, split large workflows with more than 100 nodes:
Split by business domain: Split processing pipelines for different business topics (such as transactions, users, and products) into separate workflows.
Use SUB_PROCESS to encapsulate common logic: Encapsulate common, reusable processing steps (such as data cleansing and formatting) as referenceable workflows.
Scheduling dependency settings
Scheduling dependencies connect isolated tasks into an orderly data production pipeline through two conditions: scheduled time fulfillment and upstream success. When you orchestrate nodes within a workflow, scheduling dependencies are automatically established. You can also configure more complex dependencies through scheduling dependency settings.
For more information, see Configure scheduling dependencies.
Workflow-level dependencyUse workflow-level dependencies when the entire workflow must wait for other tasks (another workflow or a standalone node) to complete before starting. This suits scenarios where a workflow acts as an independent business module. For example, a sales workflow waits for the output of a base data workflow before starting. | Node-level dependencyUse node-level dependencies when a specific node in a workflow must wait for an external task outside the current workflow to complete. This enables fine-grained cross-workflow orchestration. For example, a summary node in a report workflow waits for a specific node in an external financial system to produce output. We recommend configuring a dependency check node upstream to ensure that the dependent task completes on time. |
Cross-cycle dependencyCross-cycle dependency means that the current cycle instance of a task depends on instances from a different cycle, supporting same-node self-dependency or cross-node dependency. For example, tasks that use INSERT OVERWRITE to overwrite partitions in a table, or scenarios involving cumulative calculations. After self-dependency is enabled, the current day's instance must wait for the previous day's instance to succeed before it can run. | Cross-workspace dependencyTo depend on tasks in another DataWorks workspace, uniquely identify a node by using the workspace name and the node's Output Name, Name, or ID. This suits cross-department and cross-project data collaboration. For example, a task in a marketing workspace references key data from an accounting workspace. |
Parameter design and flow
Workspace parameters are available only in DataWorks Professional Edition and later.
DataWorks Data Studio supports four levels of parameter passing and enables flexible cross-node dynamic data transfer through node context parameters. Listed from lowest to highest scope:
Node parametersWhen the same SQL code needs to process different partition data every day, use node parameters to define dates dynamically. Constants, built-in variables, and custom time expressions are supported. For more information, see Configure scheduling parameters. | Context parametersPass values dynamically to downstream nodes through upstream output parameters. Constants, variables, and upstream run results are all supported. For more information, see Configure context parameters. |
Workflow parametersWhen dozens of nodes in a workflow need to share certain business identifiers, workflow parameters apply across all nodes in the workflow and eliminate the need to modify node parameters one by one. For more information, see Workflow parameters. Note Sub-workflows within a workflow can directly reference the workflow parameters. | Workspace parametersWhen code runs in different environments, database names and resource paths typically differ. Workspace-level parameters distinguish between environments and apply to all nodes in the workspace. For example, define a workspace parameter For more information, see Configure workspace parameters. |
Schedule settings
The key difference between scheduled workflows and legacy business processes is that a workflow sets the schedule as a whole, while a business process is only a physical grouping and does not support setting a schedule as a whole.
Schedules are set at the workflow level. Internal nodes can only set a delay execution time, calculated by adding the delay to the scheduled time defined by the workflow.
Dimension | Workflow-level configuration | Internal node-level behavior |
Time attribute | Absolute time (for example, 02:00) | Relative time (delay based on the workflow scheduling time) |
Cycle attribute | Defines daily/hourly/minute/weekly/monthly/yearly cycles | Inherits the workflow cycle and cannot be modified |
Trigger logic | Physical time arrival + upstream success | Physical time + delay time + upstream success |
Debugging and running
After completing node and workflow development, debug individual nodes and run the entire workflow to verify correctness before deployment.
Single node debugging
Single node debugging verifies code logic within one node, such as validating an SQL statement, a Python script, or a Data Integration synchronization task. This mode runs only the current node without triggering upstream or downstream dependencies.
In the Debug Configuration panel on the right side of the node, configure the following parameters:
Parameter name
Description
Compute resource
Select the associated compute resource. If no compute resource is available, select Create Compute Resource from the drop-down list.
ImportantMake sure that the compute resource and the resource group are connected. For more information, see Network connectivity.
Resource group
Select a resource group that has passed the connectivity test when the compute resource was associated. Some nodes support configuring dependency packages on the resource group to extend the runtime environment.
(Optional) Dataset
Some nodes (such as Shell and Python) support mounting datasets to access unstructured data stored in OSS or NAS.
(Optional) Script parameters
When you configure node content and define variables by using the
${parameter_name}format, you need to configure Parameter name and Parameter Value in Script Parameters. At runtime, the variables are dynamically replaced with actual values. For more information, see Configure scheduling parameters.(Optional) Associated role
Some nodes (such as Shell and Python) support configuring associated roles to access resources of other Alibaba Cloud services.
In the toolbar above the node, click Save and then Run the node task.
View the runtime log and run results at the bottom of the node.
Workflow debugging
Workflow debugging verifies whether data dependencies, parameter passing, and execution order across multiple nodes in the pipeline are correct. After single node debugging, you can debug a partial pipeline or the entire workflow.
On the DAG canvas page of the workflow, click the Run button in the top toolbar. You can also select a node, right-click it, and choose Run to this node or Run from this node to perform partial verification.
In the dialog box that appears, assign temporary values for all node variables in the workflow (such as
${bizdate}) for this debug run.The system strictly executes nodes from top to bottom according to the dependency relationships defined in the DAG. You can monitor node status in real time on the canvas and click any node to view the runtime log.
You can click the Return button on the left side of the canvas to return to the workflow development state.
The run history on the right side displays all debug run records of the workflow.
All common nodes that involve flow control (do-while, branch, and join) must be placed within a workflow and work together with upstream and downstream nodes to be effectively debugged and executed.
Management and operations
Manage workflow nodes
DataWorks allows you to import existing standalone nodes into a workflow, remove nodes from a workflow, or move nodes between workflows for efficient reuse and modular management.
Import existing nodes into a workflow
You can add existing standalone nodes that do not belong to any workflow to the current workflow canvas through Import Node for reuse
Double-click the target workflow to open the canvas editing page.
In the left component panel, switch to the Import Nodetab.
The panel lists all standalone nodes that can be added. You can filter and search by Node Type, Path, or Node Name to quickly locate the target node.
After you find the node, drag it onto the canvas to complete the import.
Remove nodes from a workflow
You can remove a node from a workflow to make it a standalone node, or move it directly to another workflow.
Remove from a workflow as a standalone node
Use this feature when you need to decouple a node from a workflow so that it is no longer part of the workflow.
In the project directory tree or on the workflow canvas, right-click the target node.
In the context menu, select Remove from workflow.
In the confirmation dialog, select the target path where the node will be stored after removal, and confirm.
Move to another workflow
Use this feature when you need to restructure a business process and migrate a node from the current workflow to another workflow.
In the project directory tree or on the workflow canvas, right-click the target node.
In the context menu, select Move to another workflow.
In the list that appears, select the target workflow and confirm.
You can also select and drag a node directly to the target workflow to move it.
Before performing this operation, carefully evaluate the potential impact on existing business processes and promptly reconfigure dependencies at the new location.
When you remove or move a node from a workflow, the upstream and downstream dependencies configured for that node within the original workflow are broken.
Workflow parameters configured on the node also become ineffective.
Clone a workflow
Cloning copies an existing workflow — including all its internal nodes, code, and dependencies — to generate an independent workflow. To clone, go to Project Directory, right-click the target workflow, and select Cloning.
Cloning copies nearly all configurations, which can introduce risks. Before deploying the cloned workflow, verify the following ("environment isolation" check):
Output tables and target data sources: The cloned workflow code writes to the same target tables as the original workflow. You must modify the target table names in the code (for example, change
ods_user_tabletodev_ods_user_table) or modify the node output configuration to prevent multiple tasks from operating on the same table, which would cause data conflicts.Upstream dependencies: Check the upstream dependency configuration of the workflow. The cloned workflow still depends on the original upstream tasks by default. Confirm whether this meets your expectations. Otherwise, it may cause data pipeline confusion or tasks to dry run.
Parameter configuration: Check the custom parameters of the workflow and its internal nodes, especially parameters related to date partitions (such as
${bizdate}), input/output paths, and other parameters. Make sure they are correct in the new environment.
In addition to cloning workflows, you can also clone nodes within a workflow, but cross-workflow node copying is not supported.
Version management
When an internal node is deployed individually, a new version is also generated for the workflow.
Version management automatically records every workflow change, supports viewing and comparing historical versions, and enables rollback to any historical state when necessary.
On the right side of the workflow canvas, the Version panel shows two types of records:
Development record: Generated each time you click Save on the canvas. This snapshot prevents accidental code loss during development and does not affect tasks running in production.
Deployment record: Generated after a workflow is deployed to the production environment. This is the version that the production scheduler executes. When production tasks encounter issues, focus on deployment records for rollback.
Typical use cases
Quick rollback on failure: When a production task fails due to code changes, use Restore to revert the workflow to the last stable deployment version with one click to minimize downtime.
Change audit and tracing: To investigate when logic was changed and by whom, use the version list and Diff features to trace each code change.
Code recovery: If you accidentally delete code logic without saving, restore it from the most recent development record.
Important considerations
Restore overwrites the current development area: Restore uses the selected historical version to overwrite all code and configurations on the current canvas. Before restoring, back up any uncommitted changes by copying them to a local text editor.
Redeployment is required after a restore: Restore only restores code to the development state. You must manually deploy it for the rollback to take effect in production.
Deployment and operations
Node/workflow deployment: After you complete workflow development, deploy the nodes/workflows to the production environment. At this point, nodes in the development environment generate corresponding scheduled tasks in the production environment. For more information, see Deploy nodes and workflows.
Node/workflow operations: Workflows in the production environment are periodically scheduled based on schedule settings. Go to Operation and Maintenance Center to view the scheduling status of scheduled workflows and perform related operations. For more information, see Scheduled tasks and Scheduled instances.
Task execution mechanism
In scheduling scenarios, the overall success of a workflow depends on the run status of its internal tasks.
A workflow's overall success is determined by the final status of its internal nodes:
Failure: If any critical node fails, the entire workflow instance is typically marked as failed.
Freeze/Pause: If a node within the workflow is manually frozen or paused, and that node is an upstream node for subsequent nodes, the pipeline is interrupted and the entire workflow is also marked as failed.
Join node: If the workflow uses a join node, even if a node in an upstream branch fails, the entire workflow may return "success" as long as the join node's own logic determines success. Therefore, carefully design the upstream checks for join nodes to prevent tasks from running with errors undetected.
Special scenarios:
When you freeze a backfill data instance of a workflow task, the workflow instance is set to success status.
In backfill data scenarios, if the system determines that a task cannot be executed, the workflow is set to failed status.
There is a latency between the instance status update and the actual occurrence of the failure event.
Quotas and limits
Node count limit: A single workflow supports up to 400 internal nodes. To ensure canvas loading performance and maintainability, keep the node count under 100. For larger scenarios, use modular splitting.
Sub-workflow limits:
A workflow with the referenceable option enabled cannot have its nodes depend on any external tasks, nor can it be directly depended on by any external tasks. Otherwise, an error occurs during deployment.
After such a workflow is deployed to the production environment, scheduled instances are not automatically generated by default. The workflow runs as a subtask only when referenced by a SUB_PROCESS node in another workflow.
Maximum parallel instances: Scheduled workflows do not support setting the maximum number of parallel instances at the workflow level. You can only set this for internal tasks within the workflow. To limit concurrent execution, configure Max Parallel Instances in the scheduling policy of individual nodes. For more information, see Configure scheduling policies.
Unsupported node types: EMR Spark Streaming, Flink SQL Streaming, Flink JAR Streaming, and Flink Python Streaming are not supported in workflows. They can only be developed and run as standalone nodes.
FAQ
Q: What is the difference between a workflow and a business process?
A: A workflow is a unified scheduling entity, while a business process is simply a folder-based grouping.
Q: Why does debugging succeed but periodic scheduling fails?
A: The most common cause is environment inconsistency. Focus on the following:
Resource group differences: During debugging, you may have used a personal or debug resource group, while production scheduling uses a production resource group. Verify that the production resource group is valid, has sufficient capacity, and has the correct permissions.
Permission differences: The production environment execution account may lack access permissions to certain tables, functions, or resources.
Dependency differences: Dependencies in the production environment are inconsistent with those in the development environment, or the output of upstream dependencies does not exist in the production environment.
Q: Why is an instance always pending/not running?
A: An instance runs only when all conditions are met: scheduled time, available scheduling resources, and upstream dependency completion.
Upstream dependencies not completed: In Operation Center, view the dependency view of the instance to check which upstream task has not yet succeeded.
Resource wait: The compute resource group queue is full, and the task is waiting in line for resources.
Workflow/node frozen: Check whether the workflow or its upstream nodes have been manually frozen or paused.
Instance not generated: Verify that the current time has reached the workflow's scheduled time. If the task was just deployed, you need to wait until the next schedule cycle for instances to be generated. For more information, see Scheduling time and instance generation.
Q: Why does the entire workflow show as failed when some nodes succeeded?
A: This is determined by the workflow status evaluation rules:
Critical path failure: If any node with incomplete downstream dependencies fails, the entire workflow is marked as failed, even if other parallel branches have succeeded.
Node frozen: An internal node being frozen causes the entire workflow to fail.