Data Quality lets you configure quality monitoring rules for data tables. You can use these rules to monitor whether table data meets your requirements, automatically block problematic tasks, and prevent dirty data from propagating downstream, ensuring the output data meets expectations. This topic describes how to configure, run, and manage quality monitoring rules for a specific table.
Prerequisites
Before you can configure quality rules for an engine's data tables, you must first acquire its metadata. For more information, see Metadata acquisition.
Limitations
-
Data source limitations: Quality monitoring rules are supported only for MaxCompute, E-MapReduce, Hologres, CDH Hive, AnalyticDB for PostgreSQL, AnalyticDB for MySQL, StarRocks, MySQL, SQL Server, DLF, and Lindorm data sources.
-
Network limitations: After you configure a rule, the scheduling node that generates the table data must use a resource group with an active network connection to trigger the Data Quality rule check.
-
Rule effectiveness limitations: Rules that use a dynamic threshold require at least 21 days of sample data to function correctly. If you have fewer than 21 days of data, the rule check may fail or produce inaccurate results. In this case, you can configure the rule, associate it with a scheduling task, and then use the backfill feature to generate the required sample data.
Core components of quality monitoring
Configuring quality rules for a table creates a complete quality monitoring plan. This plan consists of four key components:
-
Monitoring scope: Specifies the target asset for data quality checks. The configuration includes:
-
Monitored Object: The physical table or tables selected for data quality checks. You can monitor both partitioned and non-partitioned tables.
-
Data range: For a partitioned table, a partition filter expression dynamically defines which partitions to scan during each check. For example, use
$[yyyymmdd-1]to check the partition data from the day before the business date.
-
-
Monitoring Rule: Defines the specific validation logic and measurement standards to determine whether data meets expectations.
-
Rule definition: You can add one or more quality rules to a monitored object. Each rule is instantiated from a rule template. The template can be one of the following types:
-
System template: A built-in DataWorks template that covers multiple dimensions, such as integrity, uniqueness, and validity. Examples include "table row count fluctuation" and "field unique value count".
-
Custom template: A user-defined template for creating reusable and personalized validation logic by writing SQL.
-
-
Rule properties: You must configure key properties for each rule, including a threshold (for example, a fluctuation rate not exceeding 30%) and its severity (strong rule or weak rule). If a check for a strong rule fails, it can block the associated scheduling task.
-
-
Trigger Method: Defines when the quality monitoring task runs.
-
Scheduled trigger: Associates the quality monitoring with an upstream DataWorks scheduling node, typically the one that generates the monitored table. When the scheduling node runs successfully, it automatically triggers the associated quality rules. This is a best practice for automated data quality assurance.
-
Triggered Manually: The validation process is not associated with a scheduling task; you must start it manually from the UI. This method is suitable for temporary, one-time data exploration and validation.
-
-
Alert policy: Configures the notification strategy for when data quality issues are found.
-
Alert subscription: You can configure alerts for specific rule check results, such as "failed" or "warning". The system supports sending notifications through various channels, including email, SMS, phone calls, DingTalk, Lark, WeCom chatbots, and custom webhooks.
-
Configuring these four components and saving the settings creates a complete quality monitoring plan. Before you deploy it to the production environment, we recommend that you use the test run feature to verify your configuration.
Procedure
Step 1: Go to the table quality details page
-
Log on to the DataWorks console. In the target region, click in the left-side navigation pane. Select a workspace from the drop-down list and click Go to Data Quality.
-
Go to the rule configuration page for tables.
In the navigation pane on the left, click to go to the rule configuration page.
-
From the Connection list on the left, select the database that contains the table for which you want to configure rules.
-
Filter tables by criteria such as database type, database, and table name. Click the target table name, or click Rule Management in the Actions column to open the table's quality details page.
This page displays all quality monitors and rules for the table. You can filter rules by whether they are associated with a quality monitor and define the run mode for unassociated rules.
-
Step 2: Create a quality monitor
-
Create a quality monitor.
You can create a quality monitor in one of the following ways:
-
Method 1: On the Table Quality Details page, click the Rule Management tab. Click the
icon next to Monitor Perspective to create a quality monitor. -
Method 2: On the Table Quality Details page, switch to the quality monitoring tab. Click Create Monitor.
-
-
Configure the parameters for the quality monitor.
Parameter
Parameter
Description
Basic Configurations
Monitor Name
Enter a custom name for the quality monitor.
Quality Monitoring Owner
Specify an owner for the quality monitor. In an alert subscription, you can set this owner as the recipient for Email, Email and SMS, or Telephone notifications.
Monitored Object
The object to be checked for data quality. By default, this is the current table.
Data Scope
Use a partition filter expression to define the table partitions to be checked.
-
Non-partitioned table: You do not need to configure this parameter. By default, the Full Table is checked.
-
Partitioned table: The expression must be in the
partition_name=partition_valueformat. The partition value can be a fixed value or a value from Appendix 2: Built-in partition filter expressions.
NoteThis setting is ignored when you configure rules by using a custom template or custom SQL, as the custom SQL itself determines which partitions are checked.
Select Monitoring Rule
Select Monitoring Rule
Select the data quality rules to apply to the specified data range.
Note-
You can create multiple quality monitors for different partitions and associate each with different data quality rules.
-
If you have not created a data quality rule, you can skip this step. You can create the quality monitor first and then add rules to it later. For more information about how to create a data quality rule, see Step 3. Configure data quality rules.
Running Settings
Trigger Method
The method that triggers the quality monitor.
-
Triggered by Node Scheduling in Production Environment: Associates the quality monitor with a specified periodically scheduled task in DataWorks Operation Center. After the task runs successfully, the data quality rules in this quality monitor are automatically triggered. Dry-run tasks do not trigger data quality rule checks.
-
Triggered Manually: Manually triggers the data quality rules that are associated with the current quality monitor.
ImportantIf the table that you monitor is not a MaxCompute table and you set the Trigger Method to Triggered by Node Scheduling in Production Environment, the selected periodically scheduled task cannot use the public scheduling resource group. Otherwise, an error occurs when the quality monitor runs.
Associated Auto Triggered Node
If you set the Trigger Method to Triggered by Node Scheduling in Production Environment, you can configure this parameter to specify an associated scheduling node. After the specified scheduling node runs successfully, the data quality rules are automatically triggered.
Resources
The computing resources used to run the data quality rule checks. By default, the data source of the monitored table in the workspace is selected. If you select another data source, make sure that its resources can access the table.
Handling Policies
Quality Issue Handling Policies
The policy to apply when a data quality issue is detected.
-
Alert: When a data quality issue is detected, an alert notification is sent to the quality monitor's subscribers.
The default conditions are:
Strong rule - Critical anomaly,Strong rule - Warning anomaly,Strong rule - Check failed,Weak rule - Critical anomaly,Weak rule - Warning anomaly, andWeak rule - Check failed. -
Blocks: When a data quality issue is detected, the system fails the triggering production scheduling node and blocks its downstream nodes. This action blocks the production pipeline to prevent problematic data from spreading.
The default condition is
Strong rule - Critical anomaly.ImportantIf you set the policy to Blocks, an alert is also triggered when a data quality issue is detected.
Alert Method Configuration
You can send alert notifications via Email, Email and SMS, DingTalk Chatbot, DingTalk Chatbot @ALL, Lark Group Chatbot, Enterprise WeChat Chatbot, Custom Webhook, or Telephone.
Note-
To use a DingTalk, Lark, or Enterprise WeChat chatbot, add the chatbot to obtain its webhook URL, and then paste the URL into the alert subscription settings.
-
The Custom Webhook method is supported only in DataWorks Enterprise Edition. For information about the message format of an alert notification that DataWorks sends via a Custom Webhook, see Appendix: Webhook message format.
-
If you select Email, Email and SMS, or Telephone as the notification method, you can set Authorized object to Data Quality Monitoring Owner, Shift Schedule, or Scheduling Task Owner.
-
Data Quality Monitoring Owner: Sends alert notifications to the Quality Monitoring Owner specified in the Basic Configurations section of the current quality monitor.
-
Shift Schedule: When an associated scheduling node triggers a quality alert, the system sends a notification to the on-duty user for the current day in the shift schedule.
-
Scheduling Task Owner: Alert notifications will be sent to the Head of the scheduling node associated with the quality monitor.
-
-
-
Click Save to create the quality monitor.
Step 3: Configure data quality rules
You can configure quality rules based on built-in table-level and field-level rule templates. For more information about built-in rule templates, see View built-in rule templates.
-
On the Table Quality Details page, click the Rule Management tab, select the quality monitor that you created, and then click Create Rule to go to the rule configuration page.
-
Create a data quality rule.
Data quality provides the following methods to configure rules. Select the method that best fits your business requirements.
Method 1: System template
Data quality provides dozens of built-in rule templates. In the pane on the left, click + Use next to a template to quickly create a quality rule. You can add multiple rules at the same time.
You can click + System template rule at the top and then modify the Template parameter to change the rule template.
Method 2: Custom template
NoteBefore you can create a rule by using a custom template, you must go to to create a custom rule template. For more information, see Create and manage custom rule templates.
When you use a custom template, its basic configurations, such as the FLAG parameter and validation SQL, are automatically populated. You can specify a custom Rule Name and configure monitoring thresholds based on the rule type. For example, a numeric rule requires a normal threshold and a critical threshold, while a fluctuation-type rule also requires a warning threshold.
Method 3: Custom SQL
This method allows you to customize the data quality validation logic for the table.
Method 4: Custom script
Custom script rules support hour-level and minute-level data validation. For information about how to write a script rule, see Use a system rule template. Example:
- assertion: change 30 minutes ago for max(id) = 15 name: 30-minute difference in max value of id field is 15 -
(Optional) Add the configured rule to a quality monitor. For more information about quality monitors, see Step 2. Create a quality monitor.
NoteA quality rule can be triggered only after you add it to a quality monitor. You can select an existing quality monitor here, or you can select this quality rule in the Select quality rules step when you configure a quality monitor.
-
Click Confirm.
Step 4: Test the rule
You can test a quality monitor's rules in the following ways.
Rule management tab
-
On the Rule Management tab, under Monitor Perspective, find the quality monitor that you created and click Test Run.
-
In the Test Run dialog box, confirm parameters such as Data Scope and Scheduled Time, and then click Test Run. When the Started message appears, you can click View Details to view the test run details.
Quality monitoring tab
-
On the Monitor tab, find the quality monitor that you created, and click Test in the Actions column.
-
In the Test Run dialog box, confirm parameters such as Data Scope and Scheduled Time, and then click Test Run. When the Started message appears, you can click View Details to view the test run details.
Step 5: Modify alert subscriptions
You configured an alert subscription in Step 2. Create a quality monitor. When a rule is triggered, the system sends a notification to the specified alert recipients. If you need to change the alert recipients, you can modify the alert subscription in one of the following ways.
Rule management tab
-
On the Rule Management tab, under Monitor Perspective, find the quality monitor that you created, click
, and then select Alert Subscription. -
In the Alert Subscription dialog box, after you add a Notification Method and a Recipient, click Save in the Actions column. After saving, you can add more notification methods.
The supported notification methods include Email, Email and SMS, DingTalk Chatbot, DingTalk Chatbot @ALL, Lark Group Chatbot, Enterprise WeChat Chatbot, Custom Webhook, and Telephone.
Note-
To use a DingTalk, Lark, or Enterprise WeChat chatbot, add the chatbot to obtain its webhook URL, and then paste the URL into the alert subscription settings.
-
The Custom Webhook method is supported only in DataWorks Enterprise Edition. For information about the message format of an alert notification that DataWorks sends via a Custom Webhook, see Appendix: Webhook message format.
-
If you select Email, Email and SMS, or Telephone as the notification method, you can set Authorized object to Data Quality Monitoring Owner, Shift Schedule, or Scheduling Task Owner.
-
Data Quality Monitoring Owner: Sends alert notifications to the Quality Monitoring Owner specified in the Basic Configurations section of the current quality monitor.
-
Shift Schedule: When an associated scheduling node triggers a quality alert, the system sends a notification to the on-duty user for the current day in the shift schedule.
-
Scheduling Task Owner: Alert notifications will be sent to the Head of the scheduling node associated with the quality monitor.
-
-
Quality monitoring tab
-
On the Monitor tab, find the quality monitor that you created, and choose in the Actions column.
-
In the Alert Subscription dialog box, after you add a Notification Method and a Recipient, click Save in the Actions column. After saving, you can add more notification methods.
The supported notification methods include Email, Email and SMS, DingTalk Chatbot, DingTalk Chatbot @ALL, Lark Group Chatbot, Enterprise WeChat Chatbot, Custom Webhook, and Telephone.
Note-
To use a DingTalk, Lark, or Enterprise WeChat chatbot, add the chatbot to obtain its webhook URL, and then paste the URL into the alert subscription settings.
-
The Custom Webhook method is supported only in DataWorks Enterprise Edition. For information about the message format of an alert notification that DataWorks sends via a Custom Webhook, see Appendix: Webhook message format.
-
If you select Email, Email and SMS, or Telephone as the notification method, you can set Authorized object to Data Quality Monitoring Owner, Shift Schedule, or Scheduling Task Owner.
-
Data Quality Monitoring Owner: Sends alert notifications to the Quality Monitoring Owner specified in the Basic Configurations section of the current quality monitor.
-
Shift Schedule: When an associated scheduling node triggers a quality alert, the system sends a notification to the on-duty user for the current day in the shift schedule.
-
Scheduling Task Owner: Alert notifications will be sent to the Head of the scheduling node associated with the quality monitor.
-
-
Next steps
After a quality monitor runs, go to Quality O&M in the navigation pane on the left and click Monitor and Running Records to view the table's quality check status and the complete records of its quality rule checks.
Appendix
Appendix 1: Fluctuation rate and variance formulas
-
Formula for fluctuation rate:
fluctuation rate = (sample value - baseline value) / baseline value-
sample value: The value of the current day's sample. For example, for a 1-day fluctuation rate check on the table row count of an SQL task, the sample value is the row count of the current day's partition.
-
baseline value: A reference value from historical samples.
Note-
If the rule is an SQL task
table row count, 1-day fluctuation ratecheck, the baseline value is the row count of the previous day's partition. -
If the rule is an SQL task
table row count, 7-day average fluctuation ratecheck, the baseline value is the average row count from the previous 7 days.
-
-
Formula for variance:
(current sample value - average of last N days) / standard deviationNoteYou can use variance only for numeric types such as BIGINT and DOUBLE.
Appendix 2: Built-in partition expressions
Assume the following scenario:
-
The data timestamp (bizdate) is
20240524 -
The scheduling time is
10:30:00
|
Partition expression |
Description |
Example |
|
|
Checks the partition data for the current data timestamp. |
|
|
|
Checks the partition data for the day before the data timestamp. |
|
|
|
Checks the partition data for 7 days before the data timestamp (one week ago). |
|
|
|
Checks the partition data for the same day of the previous month. |
|
|
|
Checks the partition for the current data timestamp, down to the second of the scheduling time. |
|
|
|
Checks the partition for midnight (00:00:00) of the current data timestamp. |
|
|
|
Checks the partition for the time one hour before the scheduling time. |
|
|
|
(For an hourly partition) Checks the partition for the previous hour. The format is usually |
|
|
|
(For a minute-level partition) Checks the partition for the time 30 minutes before the scheduling time. The format is usually |
|
|
|
(For a two-level partition) Checks all hourly partitions for the day before the data timestamp. |
All partitions from |