All Products
Search
Document Center

DataWorks:Configure rules by template

Last Updated:Aug 24, 2026

Data Quality provides a variety of preset table-level and field-level monitoring templates. This topic describes how to configure monitoring rules using a template.

Limitations

Templates can be used to configure monitoring rules for the following data sources: MaxCompute, EMR, Hologres, CDH Hive, AnalyticDB for PostgreSQL, AnalyticDB for MySQL, StarRocks, MySQL, Lindorm, SQL Server, and DLF.

For a discrete-value rule template, you can specify only one field for grouping. If you specify multiple fields, only the first field takes effect. To perform grouped checks on multiple fields, create a separate rule for each field.

Procedure

Follow these steps to configure quality rules by using a template:

  1. Select a rule template and configure the check method.

    There are two types of built-in rule templates: table-level and field-level. After you select a template, you define the check method for a quality rule that targets the table. This quality rule specifies how to check the table data to verify that it meets expectations.

  2. Add tables or fields to check in batches

    Select the tables or fields to check in batches, and apply the rule template to them.

  3. Create or associate a quality monitor

    Associate quality rules with a quality monitor to define the quality checks for a specific object: a Data Scope of a table (such as a specific partition of a partitioned table).

Procedure

Step 1: Go to the Configure by Template page

  1. Log on to the DataWorks console. In the target region, click Data Governance > Data Quality in the left-side navigation pane. Select a workspace from the drop-down list and click Go to Data Quality.

  2. In the left-side navigation pane, choose Rule Setting > Configure by Template to open the Configure by Template page.

    Data Quality provides built-in Table-level and Field Level rule templates. Click View Monitoring Rules for a template to configure monitoring rules for tables or fields in batches.

    The Configure by Template page provides three filters at the top: Keyword Search, Associated Scope, and Quality Dimension. The built-in table-level templates include various timeliness monitoring templates, such as fixed value of table rows, 1-day difference in table rows, 1-day fluctuation rate of table rows, 7-day fluctuation rate of table rows, and 30-day fluctuation rate of table rows. The template description explains the baseline calculation method and comparison logic for each template.

Step 2: Configure monitoring rule properties

  1. Select a template to apply to multiple tables or partitions in batches. In the Actions column, click View Monitoring Rules. The Batch Add Monitoring Rules page for that template opens.

  2. Configure the Basic Properties of the monitoring rule.

    Parameter

    Description

    Connection Type

    The data source type for the tables that this rule applies to.

    Note

    Templates can be used to configure monitoring rules for the following data sources: MaxCompute, EMR, Hologres, CDH Hive, AnalyticDB for PostgreSQL, AnalyticDB for MySQL, StarRocks, MySQL, Lindorm, SQL Server, and DLF.

    Rule Source

    This value is fixed to Built-in Template and cannot be changed. It reflects the rule template you selected. For more information about built-in rule templates, see View built-in rule templates.

    Template

    Rule Name

    The system automatically generates a rule name. You can customize the suffix.

  3. Configure the advanced properties of the monitoring rule.

    Parameter

    Description

    Degree of importance

    The severity of the rule.

    • Strong rule: An important rule. By default, a critical anomaly blocks the associated scheduling task.

    • Weak rule: A regular rule. By default, a critical anomaly does not block the associated scheduling task.

    Comparison Method

    Defines how the rule validates that the table data meets your expectations.

    • Manual Settings: Customize how the data output is compared with the rule.

      The available comparison methods vary based on the rule template. The methods displayed on the UI prevail.

      • Supports Numeric Type result comparison, which is typically compared against a fixed value (expected value). Comparison methods include Greater Than, Greater Than or Equal To, Equal To, Unequal To, Less Than, and Less Than or Equal To. You can customize the normal data range (normal threshold) and abnormal data range (red threshold).

      • Supports Fluctuation result comparison, which is typically a range comparison. Comparison methods include Absolute Value, Raise, and Drop. You can customize the normal data range (normal threshold). You can also define an anomaly (orange threshold) and an unexpected result (red threshold) based on the degree of deviation.

    • Intelligent Dynamic Threshold: You do not need to manually configure fluctuation thresholds or expected values. The system uses an intelligent algorithm to automatically determine a reasonable threshold. If a data anomaly is detected, an alert is immediately triggered or the task is blocked. Dynamic thresholds also support strong and weak rules.

      Note

      The intelligent dynamic threshold comparison method is supported only for Custom SQL, Custom Scope, and Dynamic Threshold quality rules.

    Monitoring Threshold

    • If Comparison Method is set to Manual Settings, you can set the Normal threshold and Error Threshold.

      • Normal threshold: If the check result meets the value that you specify, the data check passes.

      • Error Threshold: If the check result meets the value that you specify, the data check fails.

    • If the rule performs a Fluctuation check, you must specify the Warning Threshold.

      • Warning Threshold: If the check result meets the value that you specify, the data is anomalous but does not affect business operations.

    Status

    The rule's status, which can be Enable or Deactivate. This setting controls whether the rule runs in the production environment.

    Important

    If you set the status to Deactivate, the rule cannot be triggered for a test run or by an associated scheduling task.

  4. Click Next to go to the Generate Monitoring Rule page.

Step 3: Add tables or fields

Depending on the table-level rule template or field-level rule template you select, batch add the tables or fields to which you want to apply the rule.

Tables

  1. Click Add Table. In the Batch Create dialog box that appears, select the tables for which you want to configure rules.

    Note

    The list displays all tables that match the Connection Type configured in the Basic Properties section in the previous step. You can also enter a Table Name to filter the results.

  2. After you select the tables for which you want to configure monitoring rules, click Confirm to add them to the Tables to Be Configured list.

Fields

  1. Click Add Fields. In the Select a field dialog box, select the table that contains the field you want to monitor.

    Note

    The Tables to Be Selected area displays all tables for the Connection Type that you configured in the Basic Properties section in the previous step.

  2. After you select a table, the Select a field area lists all of its fields. You can filter the fields by Field Name and Field Description.

    For example, in the Select Table pane on the left, select ods_user_info_d (User Behavior Analysis case - User Profile table). The field list for this table is displayed on the right. Select the desired fields, such as uid, dt, gender, age_range, and zodiac, and then click Add.

  3. After you select the fields for which you want to configure monitoring rules, click Create to add them to the Fields to Be Configured list.

Step 4: Create or associate a quality monitor

A quality monitor defines which rules check a specific data object. The object is a data range within the table to check, such as a specific partition of a partitioned table.

You can configure monitors individually or in batches.

Batch configuration

  1. Select one or more tables or fields to which you want to add rules, and click Configure Monitor.

  2. You can perform batch Automatically Associate, batch Disassociate, and Batch Add operations.

    • Automatically Associate: Automatically associates the selected tables or fields with existing quality monitors.

    • Disassociate: Disassociates the quality monitors from the selected tables or fields.

    • Batch Add: Configure the data range and running settings for quality monitoring on the selected tables.

      Parameter

      Description

      Data Scope

      Partitioned Table

      Defines the table partitions to check by using a partition expression.

      • Non-partitioned table: The entire table is checked by default. You can specify a range by using a WHERE clause.

      • Partitioned table: The expression format is partition_name=partition_value. The partition value can be a fixed value or a variable.

      Running Settings

      Trigger Method

      The method used to trigger the monitor.

      • Triggered by Node Scheduling in Production Environment: The quality rules under this quality monitor are automatically triggered after a specified scheduled task in Operation Center runs. A dry-run task does not trigger quality rule checks.

      • Triggered Manually: Manually triggers the quality monitoring rules associated with the current quality monitor.

      Important

      If the table that you check is not a MaxCompute table and you set Trigger Method to Triggered by Node Scheduling in Production Environment, the selected scheduled task cannot use a public scheduling resource group. Otherwise, an error occurs when the quality monitor runs.

      Associated Auto Triggered Node

      If you set Trigger Method to Triggered by Node Scheduling in Production Environment, you can use this parameter to specify an associated scheduling node. The quality monitoring rule is automatically triggered after the specified scheduling node runs.

      Resources

      The computing resources required to run the quality rule check. By default, the data source of the monitored table in the workspace is selected. If you select another data source, make sure that its resources can access the table.

Single-table configuration

  1. In the Monitor column to the right of the target table name or field, you can associate the table or field with a quality monitor. You can select an existing quality monitor or choose to Create Monitor.

  2. If no quality monitor is available, click Create Monitor. The following table describes the parameters.

    Parameter

    Parameter

    Description

    Basic Configurations

    Monitor Name

    A custom name for the quality monitor.

    Quality Monitoring Owner

    Specify an owner for the quality monitor as needed. When you configure alert subscriptions, you can specify the quality monitor owner as the alert recipient by using Email, Email and SMS, or Telephone.

    Monitored Object

    The object to be checked by Data Quality. By default, this is the current table.

    Data Scope

    Defines the table partitions to check by using a partition expression.

    • Non-partitioned table: You do not need to configure this parameter. The Full Table is checked by default.

    • Partitioned table: The expression format is partition_name=partition_value. The partition value can be a fixed value or a variable.

    Note

    This setting does not take effect for rules configured by using custom templates or custom SQL. The partitions checked by such rules are determined by the custom SQL.

    Select Monitoring Rule

    Select Monitoring Rule

    Associate quality rules with a quality monitor to determine which rules are used to check whether the data in the current data range of the table meets expectations.

    Note
    • You can create multiple quality monitors for different partitions and associate them with different quality rules to implement partition-specific checks.

    • If you have not created any quality rules, you can skip this step and create the quality monitor first. You can add rules to the quality monitor later when you create them. For more information about how to create quality rules, see Create and manage monitoring rules for a single table.

    Running Settings

    Trigger Method

    The method used to trigger the monitor.

    • Triggered by Node Scheduling in Production Environment: The quality rules under this quality monitor are automatically triggered after a specified scheduled task in Operation Center runs. A dry-run task does not trigger quality rule checks.

    • Triggered Manually: Manually triggers the quality monitoring rules associated with the current quality monitor.

    Important

    If the table that you check is not a MaxCompute table and you set Trigger Method to Triggered by Node Scheduling in Production Environment, the selected scheduled task cannot use a public scheduling resource group. Otherwise, an error occurs when the quality monitor runs.

    Associated Auto Triggered Node

    If you set Trigger Method to Triggered by Node Scheduling in Production Environment, you can use this parameter to specify an associated scheduling node. The quality monitoring rule is automatically triggered after the specified scheduling node runs.

    Resources

    The computing resources required to run the quality rule check. By default, the data source of the monitored table in the workspace is selected. If you select another data source, make sure that its resources can access the table.

    Handling Policies

    Quality Issue Handling Policies

    Configure the blocking or alerting policy for detected data quality issues.

    • Alert: When a data quality issue is detected, an alert notification is sent through the alert subscription channels of this quality monitor.

      By default, alerts are sent for the following events: strong rule·Red Anomaly, strong rule·Orange Anomaly, strong rule·Check Failed, weak rule·Red Anomaly, weak rule·Orange Anomaly, and weak rule·Check Failed.

    • Blocks: When a data quality issue is detected, the triggering production scheduling node fails. This action blocks downstream nodes and prevents problematic data from spreading.

      The default event is strong rule·Red Anomaly.

      Important

      When the policy is set to Blocks, an alert is also triggered if a data quality rule is hit.

    Alert Method Configuration

    You can send alert notifications by using methods such as Email, Email and SMS, DingTalk Chatbot, DingTalk Chatbot @ALL, Lark Group Chatbot, Enterprise WeChat Chatbot, Custom Webhook, and Telephone.

    Note
    • To add a DingTalk, Lark, or WeCom chatbot, obtain its webhook URL and paste the URL into the alert subscription settings.

    • Only DataWorks Enterprise Edition supports the Custom Webhook method. For information about the message format of Custom Webhook alert notifications that are sent by DataWorks, see Appendix: Webhook message format.

    • When you select Email, Email and SMS, or Telephone as the subscription method, you can specify the Authorized object as Data Quality Monitoring Owner, Shift Schedule, or Scheduling Task Owner.

      • Data Quality Monitoring Owner: Alert notifications are sent to the Quality Monitoring Owner specified in the Basic Configurations section of the current quality monitor.

      • Shift Schedule: When a quality rule check is triggered by an associated node, the system sends an alert notification to the on-duty personnel for the day specified in the on-duty schedule.

      • Scheduling Task Owner: Alert notifications are sent to the Head of the scheduling node associated with the quality monitor.

  3. Return to the Batch Add Monitoring Rules configuration step, click Refresh, and then select the quality monitor that you created from the Monitor drop-down list.

Step 5: Test rule execution

  1. Click Generate Monitoring Rule to go to the Verify Monitoring Rule page. On the Verify Monitoring Rule page, you can perform the following operations:

    • Test Run: Verifies that the rule configuration is valid.

      After you create the rules, you can select one or more rules and perform a Test Run. In the Test Run dialog box, select a Scheduled Time (the simulated time at which the check is triggered) and a Resource Group. The system calculates the specific partition values of the table to be checked based on this timestamp and the specified Data Scope. After the configuration is complete, click Test Run to check whether the data in the specified table partition meets the configured data quality rule.

      After a test run, you can click Running Records in the Actions column to view the details of the test run and take appropriate actions.

    • Manage Subscriptions: Defines alert recipients.

      You can send alert notifications by using methods such as Email, Email and SMS, DingTalk Chatbot, DingTalk Chatbot @ALL, Lark Group Chatbot, Enterprise WeChat Chatbot, Custom Webhook, and Telephone.

      Note
      • To add a DingTalk, Lark, or WeCom chatbot, obtain its webhook URL and paste the URL into the alert subscription settings.

      • Only DataWorks Enterprise Edition supports the Custom Webhook method. For information about the message format of Custom Webhook alert notifications that are sent by DataWorks, see Appendix: Webhook message format.

      • When you select Email, Email and SMS, or Telephone as the subscription method, you can specify the Authorized object as Data Quality Monitoring Owner, Shift Schedule, or Scheduling Task Owner.

        • Data Quality Monitoring Owner: Alert notifications are sent to the Quality Monitoring Owner specified in the Basic Configurations section of the current quality monitor.

        • Shift Schedule: When a quality rule check is triggered by an associated node, the system sends an alert notification to the on-duty personnel for the day specified in the on-duty schedule.

        • Scheduling Task Owner: Alert notifications are sent to the Head of the scheduling node associated with the quality monitor.

    • Manage Linked Nodes: Defines the trigger method for the rule.

      You can click Use Recommended Running Mode or Manually Specify Running Mode to associate one or more data quality rules with the scheduling nodes that generate the table data. In Operation Center, these nodes include periodically scheduled instances, manually triggered data backfill instances, and test instances. When a node task runs, the associated rule check is triggered. You can set the rule strength to control whether the node fails and exits, which prevents the spread of dirty data.

      • Recommended Running Mode: The system automatically associates the selected rules with the recommended scheduling nodes based on the data lineage of the nodes that produce the table.

      • Manual Running Mode: You can manually associate the selected rules with specified scheduling nodes.

      Important

      A rule must be associated with a corresponding scheduling node to be triggered automatically.

      In the Associate Scheduling dialog box: select the target Workspace and Task node, click Add, and then click OK to create the association.

    • Delete: You can select and delete one or more rules.

    • View Rule Details: In the Actions column, click View Rule Details to view a rule's details. You can also modify, enable, disable, or delete the rule, set its strength, and view logs.

  2. After the test run is successful and a schedule is associated, click Complete Check.

Next steps

After a quality monitoring run, click Monitor and Running Records under Quality O&M in the left navigation pane to view the quality check results for a specific table and the quality rule check history.

Webhook message format

This topic describes the message format and parameters of alert notifications that DataWorks sends using a Custom Webhook.

Sample message

{
  "detailUrl": "https://dqc-cn-zhangjiakou.data.aliyun.com/?defaultProjectId=3058#/jobDetail?envType=ODPS&projectName=yongxunQA_zhangbei_standard&tableName=sx_up_001&entityId=10878&taskId=16876941111958fa4ce0e0b5746379cd9bc67999d05f8&bizDate=1687536000000&executeTime=1687694111000",
  "datasourceName": "emr_test_01",
  "engineTypeName": "EMR",
  "projectName": "Online Regression Project",
  "dqcEntityQuality": {
    "entityName": "tb_auto_test",
    "actualExpression": "ds=20230625",
    "strongRuleAlarmNum": 1,
    "weakRuleAlarmNum": 0
  },
  "ruleChecks": [
    {
      "blockType": 0,
      "warningThreshold": 0.1,
      "property": "id",
      "tableName": "tb_auto_test",
      "comment": "Test rule",
      "checkResultStatus": 2,
      "templateName": "Compare the Number of Unique Field Values Against Expectation",
      "checkerName": "fulx",
      "ruleId": 123421,
      "fixedCheck": false,
      "op": "",
      "upperValue": 22200,
      "actualExpression": "ds=20230625",
      "externalId": "123112232",
      "timeCost": "10",
      "trend": "up",
      "externalType": "CWF2",
      "bizDate": 1600704000000,
      "checkResult": 2,
      "matchExpression": "ds=$[yyyymmdd]",
      "checkerType": 0,
      "projectName": "auto_test",
      "beginTime": 1600704000000,
      "dateType": "YMD",
      "criticalThreshold": "0.6",
      "isPrediction": false,
      "ruleName": "Rule Name",
      "checkerId": 7,
      "discreteCheck": true,
      "endTime": 1600704000000,
      "MethodName": "max",
      "lowerValue": 2344,
      "entityId": 12142421,
      "whereCondition": "type!='type2'",
      "expectValue": 90,
      "templateId": 5,
      "taskId": "16008552981681a0d6",
      "id": 234241453,
      "open": true,
      "referenceValue": [
        {
          "discreteProperty": "type1",
          "value": 20,
          "bizDate": "1600704000000",
          "singleCheckResult": 2,
          "threshold": 0.2
        }
      ],
      "sampleValue": [
        {
          "discreteProperty": "type2",
          "bizDate": "1600704000000",
          "value": 23
        }
      ]
    }
  ]
}

Parameters

Name

Type

Sample value

Description

projectName

String

autotest

The name of the DataWorks project.

actualExpression

String

ds=20200925

The actual partition of the data source table that was checked.

ruleChecks

Array of RuleChecks

An array of rule check results.

blockType

Integer

1

The rule strength. Valid values:

  • 1: strong rule.

  • 0: weak rule.

    You can designate important rules as strong rules based on your business requirements. If a strong rule triggers a red alert, the associated scheduling task is blocked.

warningThreshold

Float

0.1

The customizable warning threshold, which defines the acceptable deviation from the expected value.

property

String

type

The column in the data source table that the rule checks.

tableName

String

dual

The name of the table that is checked.

comment

String

Rule description.

A user-defined comment for the rule.

checkResultStatus

Integer

2

The status of the check result.

templateName

String

Compare the Number of Unique Field Values Against Expectation

The name of the monitoring template used for the check.

checkerName

String

fulx

The name of the checker.

ruleId

Long

123421

The rule ID.

fixedCheck

Boolean

false

Specifies if the check compares against a fixed value. Valid values:

  • true: The check uses a fixed value.

  • false: The check does not use a fixed value.

op

String

>

The comparison operator.

upperValue

Float

22200

The predicted upper limit, which is automatically generated after a threshold is set.

actualExpression

String

ds=20200925

The actual partition of the data source table that was checked.

externalId

String

123112232

The ID of the associated scheduling task node.

timeCost

String

10

The execution time for the check task.

trend

String

up

The trend of the check result compared to previous runs.

externalType

String

CWF2

The type of the scheduling system. Currently, only CWF is supported.

bizDate

Long

1600704000000

The data timestamp of the data being checked. For offline data, this is typically the data timestamp for the day before the check is executed.

checkResult

Integer

2

The check result.

matchExpression

String

ds=$[yyyymmdd]

The partition filter expression.

checkerType

Integer

0

The type of the checker.

projectName

String

autotest

The name of the project in the compute engine that contains the checked table.

beginTime

Long

1600704000000

The start time of the check execution, in milliseconds.

dateType

String

YMD

The type of the scheduling cycle. A common value is YMD, which represents year, month, and day.

criticalThreshold

Float

0.6

The customizable error threshold, defining the maximum acceptable deviation from the expected value. If a strong rule check exceeds the error threshold, the associated scheduling task is blocked.

isPrediction

Boolean

false

Specifies if the result is a prediction. Valid values:

  • true: The result is a prediction.

  • false: The result is not a prediction.

ruleName

String

Rule Name

The name of the rule.

checkerId

Integer

7

The ID of the checker.

discreteCheck

Boolean

true

Specifies if the check is discrete. Valid values:

  • true: The check is discrete.

  • false: The check is not discrete.

endTime

Long

1600704000000

The end time of the check execution, in milliseconds.

methodName

String

max

The method used to collect sample data. Valid values include avg, count, sum, min, max, count_distinct, user_defined, table_count, table_size, table_dt_load_count, table_dt_refuseload_count, null_value, null_value/table_count, (table_count-count_distinct)/table_count, and table_count-count_distinct.

lowerValue

Float

2344

The predicted lower limit, which is automatically generated after a threshold is set.

entityId

Long

14534343

The ID of the entity (such as a table) being monitored.

whereCondition

String

type!='type2'

The WHERE clause used to filter data for the check.

expectValue

Float

90

The expected value.

templateId

Integer

5

The ID of the monitoring template used.

taskId

String

16008552981681a0d6****

The unique ID for the check task instance.

id

Long

2231123

The primary key ID of the check record.

referenceValue

Array of ReferenceValue

An array of historical sample values for comparison.

discreteProperty

String

type1

The value of the field used for grouping if a GROUP BY clause is applied. For example, if data is grouped by a gender field, the value could be male, female, or null.

value

Float

20

The sample value.

bizDate

String

1600704000000

The data timestamp for this historical sample value.

singleCheckResult

Integer

2

The status of this individual check result.

threshold

Float

0.2

The threshold applied to this specific discrete value.

sampleValue

Array of SampleValue

An array of current sample values from the check.

discreteProperty

String

type2

The value of the field used for grouping if a GROUP BY clause is applied. For example, if data is grouped by a gender field, the value could be male, female, or null.

bizDate

String

1600704000000

The data timestamp for the current sample value.

value

Float

23

The sample value.

open

Boolean

true

Specifies if the rule is enabled.