All Products
Search
Document Center

DataWorks:Task run diagnosis

Last Updated:Aug 21, 2026

Multiple factors beyond the scheduled time that you configure in DataStudio influence a task's execution, including the scheduled times and completion times of ancestor instances, and the available resources in its resource group. This topic describes how to use the run diagnosis feature to quickly identify why a task does not run successfully.

Prerequisites

Make sure that an auto triggered instance exists. Scheduled tasks are automatically run as auto triggered instances. The Instance Generation Mode that you configure for the node in DataStudio determines how a new auto triggered instance is generated.

Background information

In Operation Center, you can determine the stage of a task run or identify the cause of a run failure based on the instance state, color, and icon. For more information, you can view the log data.

  • Instance colors and icons: Operation Center uses different colors and icons to indicate a task's stage in the run process. Each color and icon corresponds to a specific run state, as described in the following table.

  • Instance run state: You can also check why a task did not run by viewing the Node Status parameter on the Property tab of the instance details page.

No.

Status

Icon

Flowchart

1

Run Successfully

运行成功

运行流程图

2

Not run

未运行

3

Failed To Run

运行失败

4

Running

正在运行

5

Wait time

等待状态

6

Freeze

暂停冻结状态

Note

If an ancestor instance remains in the Running state for an extended period, use one of the following solutions.

  • If a non-batch synchronization node remains running and you want to find the cause, get support on DingTalk. First, click this link to join the "Alibaba Cloud Big Data & AI Platform" group. Then, scan the QR code to join the DataWorks product group, where you can ask the intelligent assistant or on-duty staff for help.技术支持二维码
  • If a batch synchronization node remains running, it may be waiting for execution resources or processing slowly. For more information, see How to troubleshoot long-running batch synchronization nodes.

Accessing the run diagnosis

If a task is not run or fails to run, find the problematic instance (such as an auto triggered instance, data backfill instance, or test instance) in Operation Center. Then, open the Intelligent Diagnosis page from the Instance Diagnose menu to quickly identify why the task did not run. The procedure is shown in the following figure.运行诊断入口

Task run diagnosis workflow

Whether a task can run successfully depends on multiple factors, such as its upstream dependencies, scheduled time, resource group, and its own execution status. If a task does not run for a long time or fails to run, use the run diagnosis feature provided by DataWorks and follow the workflow below to identify the cause.

On the Intelligent Diagnosis page, you can perform the following checks:

Check item

Description

1. Check Upstream Nodes

After you configure dependencies for a node, the node can run only after all its dependent tasks are complete. You can view the Upstream Nodes step on the Operation Details tab to locate problematic ancestor instances.

2. Check Timing Check

The node scheduling time configured in DataStudio is the expected execution time for the task. You can view the Timing Check step on the Operation Details tab to determine whether the scheduled run time has been reached.

If all ancestor instances of the current node have run successfully, which means the data it depends on has been generated, the node then checks whether its own scheduled time has been reached to decide if it should run immediately.

3. Check scheduling resources

Typically, a node is scheduled only after two conditions are met: all its ancestor instances have completed and the scheduled time of the current node has been reached.

However, because scheduling resources are limited, if the resource group for scheduling does not have enough resources to run the current task, the task waits for scheduling resources. You can view the Resources step on the Operation Details tab to check resource usage.

4. Check Execution

When the preceding conditions are met, DataWorks dispatches the task to an execution resource or service. If the task fails, you can view the Execution step on the Operation Details tab to quickly locate the cause of the failure.

5. Diagnose task alerts (Optional)

For tasks with monitoring and alerting configured, you can view the General, Influenced Baseline, and Historical instance tabs on the Intelligent Diagnostics page to see the rules or baselines that monitor the task and check their trigger status.

Upstream dependencies

This section helps you understand how an ancestor instance affects the execution of the current task and how to locate the key ancestor instances that are blocking the current task.

Impact of ancestor instances

After you configure the dependencies for a node, the node runs only after all its ancestor instances are complete. The impacts of an ancestor instance on the current task are as follows:

  • The success of an ancestor instance determines whether the current task runs

    After you configure node dependencies, a data dependency is established by default between the current task and its ancestor instances. This means that if an ancestor instance has not run, the upstream data that the current node depends on is not generated, which can cause data quality issues. Therefore, the current task must not only reach its own scheduled time but also check whether all its ancestor instances have completed.

  • The scheduled time of an ancestor instance determines the earliest start time for the current task

    An ancestor instance has its own scheduled time and will only run when that time is reached. If a downstream task's scheduled time is earlier than its ancestor's, the downstream task will not be scheduled even if its own time is reached. It must wait for the ancestor instance to complete. Therefore, the scheduled time of an ancestor instance determines the earliest possible start time for the current task. For more information, see Impact of dependencies on task execution.

Locating unexecuted ancestor instances

If a task is in the Not Run state (indicated by the 未运行 icon), you can click Run Diagnostics to go to the Operation Details tab. Then, click Upstream Nodes to automatically analyze and quickly locate the ancestor instances that have not run.

Note

The upstream analysis feature traverses up to six levels of ancestors by default to find tasks that have not run successfully. If none are found, you can click Upstream Analysis in the UI to continue the analysis.

Special cases:

  • Isolated node: If a task in the Not Run state has no ancestors when you expand the upstream nodes, it is an isolated node. An isolated node cannot be automatically scheduled. Configure its upstream dependencies.

  • Frozen ancestor: If an ancestor instance is frozen, it will block downstream tasks. Contact the owner of the ancestor instance to confirm the reason for the freeze and adjust your business logic accordingly.

Scheduled time

The node scheduling time defined in DataStudio is the expected execution time of the task. Only after all ancestor instances of the current node have run successfully (meaning the data it depends on has been generated) will the current node check if its own scheduled time has been reached. The check result determines whether to execute the task immediately. This check typically has two outcomes:

  • The scheduled time for the current task has been reached, but an ancestor instance is still running.

    In this case, as soon as all ancestor instances complete, the current task runs immediately, if sufficient resources are available in the resource group.

  • All ancestor instances have completed, but the scheduled time for the current task has not been reached.

    In this case, the task must wait until its scheduled time is reached. If the task is in the Waiting state (indicated by the 等待 icon), you can click Run Diagnostics and go to the Timing Check page for more details.

Scheduling resources

When the conditions all its ancestor instances have completed and the scheduled time of the current node has been reached are met, the current task begins to be scheduled. However, because scheduling resources are limited, if the resource group for scheduling does not have enough resources to run the current task, the task will wait for scheduling resources.

Note

Typically, scheduling resources are different from execution resources. A scheduling resource is only responsible for dispatching tasks to the execution resources of a compute engine. If a task runs for a long time and occupies engine resources, even if a scheduling resource dispatches another task to the engine, the task will be blocked due to insufficient engine resources. For more information, see the diagram in Overview of DataWorks resource groups.

Locating resource-occupying tasks

If a task is waiting for resources (indicated by the 等待 icon), you can click Run Diagnostics to go to the Resources page. There, you can see which tasks are currently occupying resources and make adjustments as needed.

Resource waiting scenarios

If a task that normally runs on schedule suddenly starts waiting for resources, check for the following scenarios.

Scenario

Description

Abnormal tasks are holding resources for a long time, blocking other tasks.

Check the Resources step on the Operation Details tab to see if any tasks are holding resources for a long time. View their run logs to determine the cause.

The number of tasks running on the resource group has increased.

An increased number of tasks on the current resource group can cause the current task to wait for resources. You can adjust the task's priority or the resource group it uses as needed.

High memory-consuming tasks exist.

Check with your team to see if any Shell or PyODPS tasks are consuming a large amount of memory from an exclusive resource group.

Important
  • The shared resource group for scheduling is shared among tenants. During peak hours (typically from 00:00~09:00), resource contention can occur, which can delay task execution. If you experience resource waiting issues with the shared resource group for scheduling, we recommend migrating your tasks to an exclusive resource group for scheduling. For more information, see Billing of exclusive resource groups for scheduling.

  • The maximum number of parallel threads supported by an exclusive resource group for scheduling depends on the specifications you purchase. For details about the number of tasks supported by different specifications, see Billing of exclusive resource groups for scheduling.

Task execution

When the preceding run conditions are met, DataWorks dispatches the task to an execution resource or service. For more information about the task dispatch mechanism, see Overview of DataWorks resource groups.

If a task is in the Failed state (indicated by the 运行失败 icon), you can go to Operation Details > Execution to view the cause of the failure.

Common causes for task execution failure include:

  • Task code execution failed (data synchronization or data processing logic failed).

  • Data quality rule validation for the output table of the task failed (the data quality rule associated with the task failed validation).

  • The task was frozen.

SQL task execution

For SQL tasks, you can view detailed execution logs directly on the Execution page within the Operation Details tab. DataWorks dispatches tasks to the corresponding engine for execution. If an SQL statement fails, consult the engine's documentation to identify the cause.

Synchronization task execution

If a Data Integration synchronization task starts to run, it means the DataWorks scheduling system has begun to schedule it. However, check the detailed execution logs to determine if the task has actually started synchronizing data. For more information on analyzing Data Integration task logs, see Analyze batch synchronization logs. Common issues with synchronization tasks are as follows:

  • Data synchronization logs print WAIT for a long time

    If the logs continuously print WAIT, it means the DataWorks scheduling system has dispatched the synchronization task, but the synchronization resource group does not have enough resources to execute it. The task is waiting for other tasks to complete and release resources.

    For example, a 4C8G exclusive resource group for Data Integration supports a maximum of eight parallel threads. If you have three tasks that are configured with three parallel threads each and two of them run concurrently, they use six threads. In this case, the machine has only two available parallel threads. The third task, which requires three threads, enters a waiting state due to insufficient resources, and the log shows wait. In this scenario, you can go to Operation Details > Execution > Data Integration to see which tasks are occupying resources and the amount of resources each task is using.

    Note
    • A Data Integration task occupies one scheduling resource. If a task does not complete for a long time, it may block other tasks.

    • If resource utilization is high but no tasks are actually running, or if tasks cannot run even though the number of executable tasks on the resource group has not reached its limit, click the application link or scan the following QR code to join the DataWorks DingTalk group for pre-sales and after-sales support. You can @ the intelligent chatbot for help or contact on-duty engineers during service hours. The QR code for the DataWorks DingTalk group is as follows.二维码

    Note

    The maximum number of parallel threads that an exclusive resource group for Data Integration can support depends on its purchased specifications. For more information, see Billing of exclusive resource groups for Data Integration.

  • Data synchronization fails

    If a synchronization task fails, refer to the detailed error message and the documentation for the specific plug-in to identify the cause. For more information, see FAQ about Data Integration.

Task alerts

For tasks with monitoring and alerting configured, click View Details in the prompt area of the Run Diagnostics page. In the Monitoring Details dialog box, view the list of rules or baselines that monitor the current task and their trigger status.

Note

This diagnostic information appears only when monitoring and alerting are configured for the task. For more information, see Diagnose task alert information.