All Products
Search
Document Center

DataWorks:O&M dashboard

Last Updated:Aug 22, 2026

The O&M dashboard provides a stability assessment, key metrics, and a resource usage overview for auto triggered nodes. It also shows operational details for manually triggered nodes and Data Integration synchronization tasks. This high-level overview of all tasks in your workspace helps you quickly find and resolve issues to improve O&M efficiency.

Usage notes

With the O&M dashboard, you can view the operations of auto triggered nodes, manually triggered nodes, and Data Integration tasks in your workspace from the following perspectives:

  • Specified workspace: View the O&M overview for a selected workspace. This view provides an O&M summary for the entire workspace and allows you to check the O&M details of Data Integration synchronization tasks separately.

  • All workspaces: View the O&M overview for all workspaces under your current account. In this view, you cannot separately check the O&M details of Data Integration synchronization tasks.

Limitations

  • The O&M dashboard is not available in the development environment of a workspace in standard mode.

    Note

    You can switch between the Production and Development environments in the top menu bar of the Operation Center interface.

  • Auto Triggered Node tab: Displays O&M data only for auto triggered nodes and their instances. Data from other types of tasks and instances is excluded.

  • Manually Triggered Node tab: Displays O&M data only for manually triggered workflows and their internal node instances.

  • Data Integration tab: Displays O&M data only for batch synchronization and real-time synchronization tasks in Data Integration.

Access the O&M dashboard

Log on to the DataWorks console. In the target region, click Data Development and O&M > Operation Center in the left-side navigation pane. Select a workspace from the drop-down list and click Go to Operation Center.

O&M for auto triggered nodes

The Auto Triggered Node tab shows the O&M overview, including the stability assessment, key issues, instance status distribution, instance completion status, and resource usage for resource groups for scheduling.

O&M stability assessment

The system assesses your workspace's operational stability based on the overall run status of its tasks.

Workspace

Single workspace

All My Workspaces

Stability chart

image

Lists all workspaces with their name, stability assessment level, number of auto triggered instances, auto triggered instance completion rate, and workspace health. The list supports paging. Click View Details on the right of a workspace to open the O&M details of that single workspace.

Description

The stability health is rated as Excellent, Good, Fair, or Poor. High-risk or low-risk labels indicate a poor health status that requires immediate attention and optimization.

  • Switch to the All My Workspaces view at the top of the page to view the O&M stability, number of auto triggered instances, and instance completion status for all workspaces you have joined.

  • You can also click View Details in the Actions column for a specific workspace to view its O&M stability details.

Key issues

This section highlights operational exceptions for auto triggered nodes from both a workspace-wide and a personal perspective, based on smart baselines and exception statistics. You can view overall issues in the workspace or only the issues related to tasks you own. This allows you to identify and resolve problems promptly to prevent business impact.

Issue type

Description

Related documentation

Illustration

Baseline instance in overtime

The number of baseline instances that are in overtime today.

An instance is in overtime if its estimated completion time exceeds the baseline's committed time, which triggers an alert.

Baseline instances

Displays eight key metrics as cards: the number of Baseline instances in overtime, Baseline instances in alert, Error events, and Slowdown events, each with its day-over-day and week-over-week change, plus the number of Isolated nodes, Frozen nodes, Expired nodes, and Nodes modified today. You can switch between the All and Mine views.

Baseline instance in alert

The number of baseline instances in an alert state today.

The alert margin helps ensure that important data is produced on time in complex dependency scenarios. If the alert margin is exceeded, nodes may fail to complete on time and exceptions may occur.

Baseline committed time and alert margin

Error event

The number of error events that occurred today.

When a task is monitored by a baseline, a run failure generates an error event. A failed task can block its downstream dependencies, so you must resolve it promptly to ensure normal operation.

Event management

Slowdown event

The number of slowdown events that occurred today.

A slowdown event occurs when a baseline-monitored task's run duration significantly increases compared to its historical average.

Isolated node

The number of auto triggered nodes that have no upstream dependencies.

An auto triggered node without upstream dependencies becomes an isolated node and will no longer run automatically.

Isolated nodes

Frozen node

The number of auto triggered nodes in the Frozen (paused) state.

When an auto triggered node is frozen, its instances are also frozen. Frozen instances do not run and will block their downstream dependencies.

Freeze and unfreeze tasks

Expired node

The number of auto triggered nodes whose scheduling validity period has ended.

A node automatically generates and runs auto triggered instances within its scheduling validity period. Outside this period, no instances are generated or scheduled.

None

Modified node

The number of auto triggered nodes modified today.

  • Modification operations: Includes changes to code, scheduling configuration, node status, and node owner.

  • Scope: Includes changes made in DataStudio and deployed to the production environment, as well as direct modifications to auto triggered nodes in the production environment.

Note

When you switch to the Mine view, this metric counts the number of modified nodes for which you are the owner.

None

Overview of auto triggered instances and nodes

The following table describes the O&M overview of auto triggered instances and nodes.

Metric

Description

Illustration

Instances Status

  • Scope: Shows the status distribution of auto triggered instances for a specific data timestamp, either for the current workspace or for instances you own. The data reflects the state at the time of the page request.

  • How to view: Click a segment of the pie chart to view the number and percentage of instances in that state.

  • Instance states to watch:

    • Failed: A failed instance may block its downstream dependent tasks.

    • Frozen: A frozen instance will not run and will block its downstream dependencies.

    • Slow-running: A running instance is considered slow-running if its duration exceeds its average duration over the past 10 days by more than 15 minutes. If there are fewer than 4 historical instances, an instance is considered slow-running if its duration exceeds 30 minutes.

Note

This metric includes only normal tasks and excludes dry-run and frozen tasks.

实例运行状态分布

Instances Completion Status

  • Scope: Shows the completion status (number of successful or not-run instances) of auto triggered instances in the current workspace between 00:00 and 23:00 on the current day, compared to the previous day and the historical average.

  • Display: The line chart shows the completion trends for today, yesterday, and the historical average. Significant deviations between the lines may indicate an anomaly that requires further investigation.

  • Node type: You can filter the results by a specific node type.

  • Historical Average: This is the average completion status over the last 10 days.

周期实例完成情况

Trend of Nodes and Instances

Scope: Shows the trend in the number of auto triggered nodes and instances in the production environment over a selected data timestamp range. You can view data for up to the last year.

周期实例与周期任务趋势

Distribution of Auto Triggered Nodes

  • Scope: Shows the number and proportion of auto triggered nodes, categorized by dimensions like node type and scheduling cycle. The data reflects the state at the time of the page request.

  • Display: The pie chart has a limit on the number of categories it can display. If the number of types exceeds the limit, they will be grouped together.

Note

In the All My Workspaces view, you can see the distribution of auto triggered nodes by workspace.

任务分布情况

Scheduling resource group usage

This section shows the utilization rate of the selected resource group for scheduling (the percentage of resources used by instances running on the group) and the trend of the number of instances running on the group over a specified time period.

Note
  • You can view data for up to seven days.

  • If resource usage exceeds 80%, we recommend scaling out the resource group to prevent resource shortages from affecting task execution.

  • Resource usage and the number of running instances are measured at the resource group level. For example, if you use an exclusive resource group for scheduling that is shared by multiple workspaces, this chart shows the total resource usage and instance count trend across all those workspaces.

调度资源组使用情况

Auto triggered instance rankings

  • Ranking of Instances on Previous Day

    This section ranks the Top 30 auto triggered instances from the previous day by run duration, resource wait time, and slow-running duration. You can use this ranking to quickly find time-consuming tasks. Click an Instance ID to go to the instance details page and view its run diagnostics.

    Note

    Slow-running duration: This is the difference between an instance's run duration yesterday and its historical average. Instances are sorted by this difference in descending order.

  • Ranking of auto triggered instances with the highest error rate in the last month

    This section ranks auto triggered instances by error rate over the last month, showing the Top 30 tasks. You can use this to quickly identify tasks with a high error rate, view their details, and find the cause of the errors.

Manually triggered nodes

The Manually Triggered Node tab shows the run status of manually triggered workflows and their internal task instances.

Manual task overview

This section shows the total number of manually triggered workflows and internal task instances that started running from a specified date, along with their success rate.

image

Run status of workflow instances

Metric

Description

Illustration

Distribution of Workflow Instances by Status

A pie chart shows the distribution of manually triggered workflow instances by run status for a given run date.

  • Click a segment to navigate to the details page for tasks in that state, where you can view and handle issues. Pay close attention to tasks that have Run failed.

  • You can view data for up to seven days.

  • When you switch to the Mine view, this chart shows the status distribution of manually triggered workflow instances for which you are the owner.

image

Workflow Ranking

Ranks workflows by run duration and failure rate for a specified run date.

  • You can use this ranking to quickly find slow or high-failure-rate workflows. Click a Node ID to open the Manual Workflow Instance details page. There, you can view the run diagnostics for a specific instance in the workflow's DAG to understand its run status.

  • Only the Top 30 workflows are displayed.

Lists the manually triggered workflow ranking with fields such as ordinal, workflow name, creator, and run duration. You can switch the sort dimension between Run Duration and Failures, and filter by run date.

Run status of internal task instances

Metric

Description

Illustration

Distribution of Internal Tasks

A pie chart shows the real-time distribution of internal task instances in Operation Center, categorized by Node Type and Owner.

image

Internal Task Ranking

Ranks internal task instances by run duration and failure rate for a specified run date.

  • You can use this ranking to quickly find slow or high-failure-rate internal task instances. Click a Node ID to go to the Manual Workflow Instance details page. On the details page, you can view the run diagnostics for a specific instance in the workflow's DAG to understand its run status.

  • Only the Top 30 internal tasks are displayed.

Lists the internal task ranking with fields such as ordinal, node ID, node name, workflow, node type, creator, and run duration. You can switch the sort dimension between Run Duration and Failures.

Data Integration tasks

The Data Integration tab shows an overview of Data Integration synchronization tasks and resource group usage for Yesterday or Today.

Data Integration resource group usage

This section shows the resource details for all Data Integration tasks in the current workspace, including the Running Tasks, Resource Usage, and Last Modified At. You can use the resource utilization and task volume to determine whether to scale resources up or down for optimal allocation. It also shows information such as the resource group name, status, number of queued tasks, specification, and region, and provides the View, Renew, and Create Resource Group entry points.

Note

Run status distribution of Data Integration synchronization tasks

A pie chart shows the distribution of synchronization tasks in the current workspace by run status. Click a segment to navigate to the details page for tasks in that state to view and handle issues. Pay close attention to Abnormal and Run failed tasks, as they often block downstream task execution.运行状态分布

Run status of batch synchronization tasks

The following table describes the run status of batch synchronization tasks.

Metric

Description

Illustration

Data Synchronization Progress

Shows the total amount of data and total traffic used for batch synchronization within the selected data timestamp.

Displays three batch synchronization progress metrics as cards: Total Data Volume, Total Public Network Traffic, and Total Records.

Statistics on amount of synchronized data

Displays data read and write curves by data source type for the selected data timestamp. This helps you quickly identify engine tasks with large data volumes and determine if they require more resources.

离线数据同步任务数据统计量

Latest Top 10 Tasks

Shows the 10 most recent Latest Failed Instance and Latest Successful Instance to give you a global view of the latest synchronization task statuses. You can use the error messages to quickly identify the cause of a failure and resolve it.

Lists the Recent Failed Instances and Recent Successful Instances with fields such as node name, node ID, start time, end time, and error message.

Execution details of synchronization tasks

You can quickly search for task instances by filtering on conditions such as Commission Time, Node Status, and Node Name to view their running details.

Lists the execution details of synchronization tasks with fields such as instance ID, node name, start time, end time, duration, source and destination data sources, synchronized data volume, number of synchronized records, and synchronization rate. You can filter by run date, node status, and node name.

Run status of real-time synchronization tasks

The following table describes the run status of real-time synchronization tasks.

Metric

Description

Illustration

Data synchronization Overview

Shows the total data rate and record rate for all real-time synchronization tasks in the current workspace.

Displays two real-time synchronization overview metrics as cards: Data Speed (BPS) and Record Speed (RPS).

Top 10 tasks with the highest latency

Shows the 10 real-time synchronization tasks with the highest latency, allowing you to quickly identify and optimize them.

Lists the real-time synchronization tasks with the highest latency, including the node name, business latency, and resource group information.

Alerts

Shows recent alert information from real-time synchronization tasks, allowing you to quickly catch and resolve exceptions.

Lists the alert records with fields such as occurrence time, node ID, node name, recipient, alert type (CRITICAL or WARNING), trigger condition, threshold, alert time, and notification method.

Failover Information

Shows failover messages for real-time synchronization tasks within a specified time frame to provide an overview of task failover status. For more information about failover, see Run and manage real-time synchronization tasks.

Lists the failover records within the specified time range with fields such as time, instance ID, node name, failover event, and actions.