The O&M dashboard provides a stability assessment, key metrics, and a resource usage overview for auto triggered nodes. It also shows operational details for manually triggered nodes and Data Integration synchronization tasks. This high-level overview of all tasks in your workspace helps you quickly find and resolve issues to improve O&M efficiency.
Usage notes
With the O&M dashboard, you can view the operations of auto triggered nodes, manually triggered nodes, and Data Integration tasks in your workspace from the following perspectives:
-
Specified workspace: View the O&M overview for a selected workspace. This view provides an O&M summary for the entire workspace and allows you to check the O&M details of Data Integration synchronization tasks separately.
-
All workspaces: View the O&M overview for all workspaces under your current account. In this view, you cannot separately check the O&M details of Data Integration synchronization tasks.
Limitations
-
The O&M dashboard is not available in the development environment of a workspace in standard mode.
NoteYou can switch between the Production and Development environments in the top menu bar of the Operation Center interface.
-
Auto Triggered Node tab: Displays O&M data only for auto triggered nodes and their instances. Data from other types of tasks and instances is excluded.
-
Manually Triggered Node tab: Displays O&M data only for manually triggered workflows and their internal node instances.
-
Data Integration tab: Displays O&M data only for batch synchronization and real-time synchronization tasks in Data Integration.
Access the O&M dashboard
Log on to the DataWorks console. In the target region, click in the left-side navigation pane. Select a workspace from the drop-down list and click Go to Operation Center.
O&M for auto triggered nodes
The Auto Triggered Node tab shows the O&M overview, including the stability assessment, key issues, instance status distribution, instance completion status, and resource usage for resource groups for scheduling.
O&M stability assessment
The system assesses your workspace's operational stability based on the overall run status of its tasks.
|
Workspace |
Single workspace |
All My Workspaces |
|
Stability chart |
|
Lists all workspaces with their name, stability assessment level, number of auto triggered instances, auto triggered instance completion rate, and workspace health. The list supports paging. Click View Details on the right of a workspace to open the O&M details of that single workspace. |
|
Description |
The stability health is rated as Excellent, Good, Fair, or Poor. High-risk or low-risk labels indicate a poor health status that requires immediate attention and optimization. |
|
Key issues
This section highlights operational exceptions for auto triggered nodes from both a workspace-wide and a personal perspective, based on smart baselines and exception statistics. You can view overall issues in the workspace or only the issues related to tasks you own. This allows you to identify and resolve problems promptly to prevent business impact.
|
Issue type |
Description |
Related documentation |
Illustration |
|
Baseline instance in overtime |
The number of baseline instances that are in overtime today. An instance is in overtime if its estimated completion time exceeds the baseline's committed time, which triggers an alert. |
Displays eight key metrics as cards: the number of Baseline instances in overtime, Baseline instances in alert, Error events, and Slowdown events, each with its day-over-day and week-over-week change, plus the number of Isolated nodes, Frozen nodes, Expired nodes, and Nodes modified today. You can switch between the All and Mine views. |
|
|
Baseline instance in alert |
The number of baseline instances in an alert state today. The alert margin helps ensure that important data is produced on time in complex dependency scenarios. If the alert margin is exceeded, nodes may fail to complete on time and exceptions may occur. |
||
|
Error event |
The number of error events that occurred today. When a task is monitored by a baseline, a run failure generates an error event. A failed task can block its downstream dependencies, so you must resolve it promptly to ensure normal operation. |
||
|
Slowdown event |
The number of slowdown events that occurred today. A slowdown event occurs when a baseline-monitored task's run duration significantly increases compared to its historical average. |
||
|
Isolated node |
The number of auto triggered nodes that have no upstream dependencies. An auto triggered node without upstream dependencies becomes an isolated node and will no longer run automatically. |
||
|
Frozen node |
The number of auto triggered nodes in the Frozen (paused) state. When an auto triggered node is frozen, its instances are also frozen. Frozen instances do not run and will block their downstream dependencies. |
||
|
Expired node |
The number of auto triggered nodes whose scheduling validity period has ended. A node automatically generates and runs auto triggered instances within its scheduling validity period. Outside this period, no instances are generated or scheduled. |
None |
|
|
Modified node |
The number of auto triggered nodes modified today.
Note
When you switch to the Mine view, this metric counts the number of modified nodes for which you are the owner. |
None |
Overview of auto triggered instances and nodes
The following table describes the O&M overview of auto triggered instances and nodes.
|
Metric |
Description |
Illustration |
|
Instances Status |
Note
This metric includes only normal tasks and excludes dry-run and frozen tasks. |
|
|
Instances Completion Status |
|
|
|
Trend of Nodes and Instances |
Scope: Shows the trend in the number of auto triggered nodes and instances in the production environment over a selected data timestamp range. You can view data for up to the last year. |
|
|
Distribution of Auto Triggered Nodes |
Note
In the All My Workspaces view, you can see the distribution of auto triggered nodes by workspace. |
|
Scheduling resource group usage
This section shows the utilization rate of the selected resource group for scheduling (the percentage of resources used by instances running on the group) and the trend of the number of instances running on the group over a specified time period.
-
You can view data for up to seven days.
-
If resource usage exceeds 80%, we recommend scaling out the resource group to prevent resource shortages from affecting task execution.
-
Resource usage and the number of running instances are measured at the resource group level. For example, if you use an exclusive resource group for scheduling that is shared by multiple workspaces, this chart shows the total resource usage and instance count trend across all those workspaces.

Auto triggered instance rankings
-
Ranking of Instances on Previous Day
This section ranks the
Top 30auto triggered instances from the previous day by run duration, resource wait time, and slow-running duration. You can use this ranking to quickly find time-consuming tasks. Click an Instance ID to go to the instance details page and view its run diagnostics.NoteSlow-running duration: This is the difference between an instance's run duration yesterday and its historical average. Instances are sorted by this difference in descending order.
-
Ranking of auto triggered instances with the highest error rate in the last month
This section ranks auto triggered instances by error rate over the last month, showing the
Top 30tasks. You can use this to quickly identify tasks with a high error rate, view their details, and find the cause of the errors.
Manually triggered nodes
The Manually Triggered Node tab shows the run status of manually triggered workflows and their internal task instances.
Manual task overview
This section shows the total number of manually triggered workflows and internal task instances that started running from a specified date, along with their success rate.

Run status of workflow instances
|
Metric |
Description |
Illustration |
|
Distribution of Workflow Instances by Status |
A pie chart shows the distribution of manually triggered workflow instances by run status for a given run date.
|
|
|
Workflow Ranking |
Ranks workflows by run duration and failure rate for a specified run date.
|
Lists the manually triggered workflow ranking with fields such as ordinal, workflow name, creator, and run duration. You can switch the sort dimension between Run Duration and Failures, and filter by run date. |
Run status of internal task instances
|
Metric |
Description |
Illustration |
|
Distribution of Internal Tasks |
A pie chart shows the real-time distribution of internal task instances in Operation Center, categorized by Node Type and Owner. |
|
|
Internal Task Ranking |
Ranks internal task instances by run duration and failure rate for a specified run date.
|
Lists the internal task ranking with fields such as ordinal, node ID, node name, workflow, node type, creator, and run duration. You can switch the sort dimension between Run Duration and Failures. |
Data Integration tasks
The Data Integration tab shows an overview of Data Integration synchronization tasks and resource group usage for Yesterday or Today.
Data Integration resource group usage
This section shows the resource details for all Data Integration tasks in the current workspace, including the Running Tasks, Resource Usage, and Last Modified At. You can use the resource utilization and task volume to determine whether to scale resources up or down for optimal allocation. It also shows information such as the resource group name, status, number of queued tasks, specification, and region, and provides the View, Renew, and Create Resource Group entry points.
-
For more information about operations on an exclusive resource group for Data Integration, see Billing of exclusive resource groups for Data Integration.
-
For more information about operations on a serverless resource group, see Use a serverless resource group.
-
The tab displays O&M data only for exclusive resource groups for Data Integration.
Run status distribution of Data Integration synchronization tasks
A pie chart shows the distribution of synchronization tasks in the current workspace by run status. Click a segment to navigate to the details page for tasks in that state to view and handle issues. Pay close attention to Abnormal and Run failed tasks, as they often block downstream task execution.
Run status of batch synchronization tasks
The following table describes the run status of batch synchronization tasks.
|
Metric |
Description |
Illustration |
|
Data Synchronization Progress |
Shows the total amount of data and total traffic used for batch synchronization within the selected data timestamp. |
Displays three batch synchronization progress metrics as cards: Total Data Volume, Total Public Network Traffic, and Total Records. |
|
Statistics on amount of synchronized data |
Displays data read and write curves by data source type for the selected data timestamp. This helps you quickly identify engine tasks with large data volumes and determine if they require more resources. |
|
|
Latest Top 10 Tasks |
Shows the 10 most recent Latest Failed Instance and Latest Successful Instance to give you a global view of the latest synchronization task statuses. You can use the error messages to quickly identify the cause of a failure and resolve it. |
Lists the Recent Failed Instances and Recent Successful Instances with fields such as node name, node ID, start time, end time, and error message. |
|
Execution details of synchronization tasks |
You can quickly search for task instances by filtering on conditions such as Commission Time, Node Status, and Node Name to view their running details. |
Lists the execution details of synchronization tasks with fields such as instance ID, node name, start time, end time, duration, source and destination data sources, synchronized data volume, number of synchronized records, and synchronization rate. You can filter by run date, node status, and node name. |
Run status of real-time synchronization tasks
The following table describes the run status of real-time synchronization tasks.
|
Metric |
Description |
Illustration |
|
Data synchronization Overview |
Shows the total data rate and record rate for all real-time synchronization tasks in the current workspace. |
Displays two real-time synchronization overview metrics as cards: Data Speed (BPS) and Record Speed (RPS). |
|
Top 10 tasks with the highest latency |
Shows the 10 real-time synchronization tasks with the highest latency, allowing you to quickly identify and optimize them. |
Lists the real-time synchronization tasks with the highest latency, including the node name, business latency, and resource group information. |
|
Alerts |
Shows recent alert information from real-time synchronization tasks, allowing you to quickly catch and resolve exceptions. |
Lists the alert records with fields such as occurrence time, node ID, node name, recipient, alert type (CRITICAL or WARNING), trigger condition, threshold, alert time, and notification method. |
|
Failover Information |
Shows |
Lists the |







