All Products
Search
Document Center

Dataphin:Create and Manage Metadata Collection Tasks

Last Updated:Aug 28, 2026

A collection task connects to a specified data source via a collection adapter, collects object metadata from the source into Dataphin, and uses built-in parsers to uniformly parse, store, and present the metadata.

Prerequisites

  • You must create an application system in Management Center > Data Source Management > Application System to serve as a collection source.

  • To collect metadata from a DataWorks application system, you must purchase the Metadata Collection value-added service.

  • By default, Dataphin supports metadata collection from relational databases. To collect metadata from other data source types, you must purchase the corresponding features.

  • For versions earlier than V5.3, you must initialize the Metadata Center configuration in the metadata warehouse tenant before collecting data from certain data sources, including AnalyticDB for MySQL 3.0, PolarDB-X (formerly DRDS), SAP HANA, and Hologres. For V5.3 and later versions, you can configure collection tasks directly without this initialization.

Limitations

  • If collected metadata objects have the same name but differ in case, the system recognizes only the format supported by the default compute engine. For example, Oracle recognizes uppercase names by default, and DM (Dameng) recognizes the object that was collected first. Other similarly named metadata objects are not processed.

  • For PolarDB-X (formerly DRDS) data sources, you can collect view objects from version 2.0 and later.

  • Due to a collection workflow upgrade, if you created a collection task for PostgreSQL, MySQL, Microsoft SQL Server, Oracle, IBM DB2, Hive (MySQL metadatabase), or StarRocks before V5.1, you cannot view historical run logs for its collection instances after upgrading to V5.1 or later until the task is run again.

  • Elasticsearch data sources do not support asset publication management.

  • When the system type is DataWorks and the collection source is MaxCompute, Dataphin collects metadata only for the projects configured in that data source.

Permissions

Super administrators, system administrators, and users granted metadata collection task management permissions through a custom global role can create and manage metadata collection tasks.

Metadata cCollection Workflow

If your data source's network is isolated from the Dataphin cluster's network, you must register a scheduling cluster. The collection task writes data to an object storage system, such as OSS, as an intermediate step before writing it to Dataphin. This process may incur additional storage costs.

Create a Collection Task

  1. In the top navigation bar of the Dataphin homepage, choose Governance > Metadata.

  2. In the left-side navigation pane, click Collection Task, and then click + New Collection Task to open the New Collection Task dialog box.

  3. In the New Collection Task dialog box, configure the parameters.

    Parameter

    Description

    Collection task name

    A globally unique name for the collection task, up to 512 characters long.

    Owner

    The owner of the collection task. You can select a member who has permissions to manage collection tasks.

    Collection task description

    A description of the collection task. The description cannot exceed 1,000 characters.

    Data origin

    Select the scope of collection sources. Metadata is collected from the selected sources. Supported data origins include data sources and application systems.

    • Data Source: Supports relational databases and big data storage databases. You can click View to go to the Data Source Management page, where the system filters for relevant data sources. For a list of supported data sources, see Supported data sources in Dataphin.

    • Application System: Select the application system from which you want to collect metadata. You can select Quick BI or DataWorks.

    Note
    • If the selected data source does not have a data source code configured, you may not be able to use the collected metadata through JDBC or in a BI platform. For information about how to configure a data source code, see Supported data sources in Dataphin.

    • You can configure only one collection task for each data source. However, you can create separate collection tasks for the development and production environments of the same data source.

    Collection scope

    Configure the collection scope based on the data source type or application system.

    • For Hive data sources, Dataphin automatically parses the corresponding dbname (database name) based on the JDBC URL configured for the data source.

    • For data sources such as MySQL, AnalyticDB for MySQL 3.0, PolarDB-X (formerly DRDS), StarRocks, OceanBase (MySQL tenant), ClickHouse, Amazon RDS for MySQL, SelectDB, Doris, DolphinDB, and TDSQL for MySQL, you can configure the collection scope by database (databases under the data source instance). You can select either All databases or Specified databases.

      • All databases: Collects from all databases that the configured account has permission to query.

      • Specified databases: Lets you specify other databases that the configured account can access. If a database is configured on the data source side, Dataphin auto-fills it by default. If you enter a database manually, the value is case-sensitive.

    • For data sources such as Oracle, PostgreSQL, Microsoft SQL Server, SAP HANA, IBM DB2, Hologres, OceanBase (Oracle tenant), Greenplum, Amazon RDS for PostgreSQL, Amazon RDS for SQL Server, Amazon RDS for Oracle, Amazon RDS for DB2, Amazon Redshift, DM (Dameng), openGauss, and TDSQL for PostgreSQL, you can configure the collection scope by schema (database names under the data source instance). You can select either All schemas or Specified schemas.

      • All schemas: Collects from all schemas that the configured account has permission to query.

      • Specified schemas: Lets you specify other schemas that the configured account can access, or quickly fill the default schema with one click. If you enter a schema manually, the value is case-sensitive.

    • When the data origin is Quick BI or DataWorks, you can configure the collection scope by workspace. You can select either All workspaces or Specified workspaces.

      • All workspaces: Collects from all workspaces that the configured account has permission to query based on the application system configuration.

      • Specified workspaces: Lets you specify other workspaces that the configured account can access based on the application system configuration.

    Note
    • For Hive and StarRocks data sources, for each partitioned table, Dataphin collects up to the 100,000 most recent partitions by creation time.

    • For OceanBase data sources, the collection scope depends on the tenant mode configured for the data source. For a MySQL tenant, metadata is collected by database. For an Oracle tenant, metadata is collected by schema.

    Data source collection scope

    This parameter is available when Data origin is set to DataWorks and you select Specified workspaces. You can configure the collection scope based on the data source by selecting either All data sources or Specified data sources.

    • All data sources: Dynamically retrieves all data sources in the specified workspace for which you have query permissions.

    • Specified data sources: Specifies other accessible data sources in the selected workspace.

    Collection object type

    This parameter is selected by default and cannot be modified. When the data origin is a data source or a DataWorks application system, the supported types are Table, View, and Field. When the data origin is a Quick BI application system, the supported type is Dashboard.

    Note
    • For an Elasticsearch data source, an index corresponds to the Table collection object type, and an index alias corresponds to the View collection object type.

    • For StarRocks data sources, collecting synchronous materialized views is not supported.

    Source system

    This parameter is available only when the data origin is a data source. Select the source system to which the collected metadata belongs. This information can be used for purposes such as filtering asset objects and displaying source system data lineage. To create a source system, see Create and manage source systems.

    Automatic data sampling

    This option is available if data sampling is enabled in Governance > Metadata > Sampling Configuration, the trigger scenario includes metadata collection, and the data source supports data preview. When enabled, the task automatically collects sample data based on the collection scope configured in Sampling Configuration > Data Source. You can modify the collection scope.

  4. Click Next to configure the collection policy.

    Parameter

    Description

    Data update policy

    New/modified metadata

    If the source system has new or updated data since the last collection, Dataphin adds new metadata and updates modified metadata. For dashboards, if a work is modified but not published (that is, its status is "Saved but not published"), Dataphin retains the previously collected published data without updating it.

    Deleted metadata

    If data has been deleted from the source system since the last collection, you can choose to Delete from Metadata List and Asset List or Ignore Deletion. For dashboards, you can choose Treat as deletion if the work status changes from "Published" to "Unpublished" or Ignore Deletion.

    • Delete from Metadata List and Asset List/Treat as deletion if the work status changes from "Published" to "Unpublished": Deletes the corresponding collected metadata. This action cannot be undone.

    • Ignore Deletion: Ignores the deletion in the source system. You can still view the object's details and history in the metadata and asset lists. You can manually delete it later.

    Data collection schedule

    Collection frequency

    Controls how often the task runs. You can select Scheduled Collection or Manual Collection.

    • Scheduled Collection: Runs the task automatically based on the configured schedule. This is suitable for scenarios that require timely metadata updates. You can set the frequency to Daily, Weekly, or Monthly, with a start time between 00:00 and 23:59. When you select a Monthly schedule, you can also choose the Last day of the month.

      If your system time zone (the time zone in your user center) differs from the scheduling time zone (the time zone configured in Management Center > System Settings > Basic Settings), the system displays both. When you configure a schedule, the system runs the task based on the scheduling time zone and displays the equivalent time in your system's time zone for your reference.

    • Manual Collection: Requires you to trigger the task manually. This is suitable for scenarios where metadata changes infrequently and you want to conserve resources.

    Runtime configuration

    Retry on error

    If a collection instance fails, you can configure it to rerun automatically based on the Number of Retries and Retry Interval.

    • Number of Retries: The maximum number of times the system automatically retries a failed instance. The default is 1. You can set this to any integer from 1 to 10.

    • Retry Interval: The time between automatic retries. The default is 5 minutes. You can set this to any value from 1 to 60 minutes.

    Note

    Retries and scheduled collections may conflict. If the next scheduled time arrives while the previous run is still in progress, Dataphin delays the next scheduled run. You can manually stop the task in the collection instance list. For more information, see View and manage collection instances.

    Execution timeout

    If a task's total runtime (excluding resource and scheduling wait times) exceeds the specified threshold, the system automatically terminates it and marks it as failed. You can set the timeout from 0 to 24 hours, with up to one decimal place.

    Scheduling resources

    The collection task consumes resources from the selected resource group's quota. To avoid high concurrency affecting other system tasks, all collection tasks created across all tenants share a unified global concurrency limit. Allocate scheduling resources carefully. You can select a resource group that is in the Normal state and belongs to the current tenant.

    The selected data source's network must be connected to the scheduling resource group's network. Otherwise, the collection task cannot run. After selecting a resource group, you can click Test Connection to test network connectivity. If the test fails, click View Logs to see the reason.

    Connection configuration

    You can view the connection configuration of the selected collection source to help you configure the collection frequency and time. For more information, see Supported data sources in Dataphin.

    Note

    The current connection configuration also applies to batch synchronization tasks, global quality monitoring rules, and metadata collection tasks.

  5. Click OK to create the collection task.

Manage Collection Tasks

  1. The collection task page displays the name, data origin and data source code, data origin type, collection method, status and time of the last collection, description, owner, effective status, task status, and last update time for each task. You can click the Data Source Management button in the upper-right corner to go to the Management Center > Data Source Management page and manage collection sources.

    You can view the Task Status of each task in the list. The available actions vary depending on the task status, as shown in the following table.

    Task status

    Actions

    Normal

    View, Edit, Ad hoc Run (for scheduled tasks), Manual Run (for manual tasks), Clone, Delete, View Metadata, View Collection Instances, and enable or disable the effective status.

    Creation Failed

    Retry, View Execution Log, View, Edit, and Delete.

    Update Failed/Deletion Failed/Activation Failed/Deactivation Failed

    Retry, View Execution Log, View, Edit, Delete, View Metadata, and View Collection Instances.

    Activating/Deactivating

    View.

    You cannot change the effective status while the task is in the Activating or Deactivating state.

    Creating/Updating/Deleting

    View.

    Abnormal

    View, Edit, Delete, View Metadata, and View Collection Instances.

  2. (Optional) You can search for a task by its name or data source name. You can also use quick filters such as My Tasks and Active Tasks, or filter by task status, effective status, owner, data origin, or collection method.

  3. In the Actions column for a task, you can perform the following operations.

    Actions

    Description

    Retry

    Reruns a failed collection task.

    View Execution Log

    Views the execution log of a failed collection task.

    View

    Views the configuration details of the collection task.

    Edit

    You cannot modify the data source type or data source. Modifying other settings does not affect the effective status.

    Ad Hoc Run

    Only scheduled collection tasks in the Normal state support ad hoc runs. If an instance from this run is still in progress when the next scheduled run time arrives, it may cause data inconsistency. If an instance of the task (either a scheduled or an ad hoc instance) is already running, you must stop the running instance before starting another.

    Manual Run

    Only manual collection tasks in the Normal state support manual runs. If an instance of the task is already running, you must stop it before starting a new run.

    Clone

    Creates a copy of the task's configuration. You must reconfigure the data source and collection scope.

    Delete

    • Delete a single task: In the Actions column, click the image icon and select Delete.

    • Delete multiple tasks: Select the checkboxes for the tasks you want to delete, and then click the image icon at the bottom of the list.

    Note

    Deleting a task does not affect any instances that are currently running. You can stop them manually if needed. After a task is deleted, no new instances are generated. You can configure a deletion policy: Delete Task and Metadata or Delete Task Only.

    • Delete Task and Metadata: Deletes the metadata collected by this task from the specified data source from both the metadata list and asset list.

    • Delete Task Only: Deletes only the collection task itself but retains the metadata that has already been collected. The metadata remains in the metadata and asset lists. If you later create a new collection task for the same data source, it may overwrite the retained metadata.

    View metadata list

    Opens the metadata list page, which is filtered to show metadata from the data source in this task.

    View collection instances

    Opens the collection instances page, which is filtered to show instances related to this task.

    Change effective status

    • Change a single task's status: In the Effective Status column, click the image switch to enable or disable the task.

    • Change multiple tasks' statuses: Select the checkboxes for the tasks you want to modify, and then click the image icon at the bottom of the list to enable or disable their effective status.

    Note

    When you activate a task, it runs automatically according to its schedule. Deactivating a task does not affect instances that are already running or pending, but it prevents new scheduled instances from being generated. You can still run the task manually.

Next Steps