Collection tasks use adapters to connect to data sources, collect object metadata from source databases into Dataphin, and parse and store it for unified display.
Prerequisites
Create an application system in Management Center > Datasource Management > Application System before you can use the application system type as a collection source.
Limits
-
If the collected metadata contains objects with the same name but different case, the system recognizes only the default format supported by the compute engine (for example, Oracle recognizes uppercase object names by default, and DM (DaMeng) recognizes objects collected first). Other same-name metadata is not processed.
-
PolarDB-X (formerly DRDS) data sources of version 2.0 or later support the collection of view objects.
-
Metadata collection for relational databases is supported by default. To collect metadata from other data source types, purchase the corresponding features.
-
Prior to version 5.3, some data sources required you to initialize the Metadata Center in the metadata warehouse tenant before collection could start. These data sources include AnalyticDB for MySQL 3.0, PolarDB-X (formerly DRDS), SAP HANA, and Hologres. In version 5.3 and later, you can configure collection tasks directly without initializing the Metadata Center.
-
Due to collection workflow upgrades, if you created collection tasks for PostgreSQL, MySQL, Microsoft SQLServer, Oracle, IBM DB2, Hive (MySQL metadatabase), StarRocks before V5.1 and upgraded to V5.1 or later without re-running them, you cannot view the historical collection instance run logs.
-
Elasticsearch data sources do not support listing management.
Permission requirements
Super administrators, system administrators, and custom global roles with metadata collection task management permissions can create and manage metadata collection tasks.
Metadata collection workflow description
If the network environment of the collected data source is not connected to the Dataphin cluster, use the register scheduling cluster feature. The collected data is first written to the object storage system (such as OSS) that Dataphin depends on as a transit, and then written to Dataphin. This process incurs additional storage costs.
Create a collection task
-
In the top navigation bar of the Dataphin homepage, choose Administration > Metadata.
-
Click Collection Task in the navigation pane on the left, and then click the +New Collection Task button to enter the New Collection Task dialog box.
-
In the New Collection Task dialog box, configure the parameters.
Parameter
Description
Collection Task Name
The name of the collection task. Must be globally unique and cannot exceed 512 characters.
Owner
The owner of the collection task. Select a member who has collection task management permissions.
Collection Task Description
Optional description for the collection task. Cannot exceed 1,000 characters.
Data Source
Select the collection source range based on the data source. Supported source types include data sources and application systems.
-
Datasource: Supports relational databases and big data storage databases. For more information, see Data sources supported by Dataphin.
-
Application System: Currently supports only Quick BI. Select the application system from which to collect metadata.
Click View to go to the Data Source Management page, where the system filters the relevant data sources for you.
Note-
If the selected data source does not have a data source encoding configured, you may not be able to use the collected metadata through JDBC or in a BI platform. For information about how to configure data source encoding, see Data sources supported by Dataphin.
-
A data source can only have one collection task. Two different environment sources (development and production) of the same data source can have separate collection tasks.
Collection Range
Configure different task collection ranges based on the data source type or application system.
-
When the data source type is Hive, the system automatically parses the corresponding dbname (database name) based on the JDBC URL configured for the data source.
-
If the data source type is MySQL, AnalyticDB for MySQL 3.0, PolarDB-X, StarRocks, OceanBase (MySQL Tenant), ClickHouse, Amazon RDS for MySQL, SelectDB, Doris, DolphinDB, or TDSQL for MySQL, you can configure the collection scope based on the database under the data source instance. You can select All Databases or Specified Database.
-
All Databases: Dynamically retrieves all databases with query permissions based on the data source configuration.
-
Specified Database: Specifies other databases with permissions based on the data source configuration. If a database is already configured for the data source, it is filled in by default. Custom database names are case-sensitive.
-
-
When the data source type is Oracle, PostgreSQL, Microsoft SQL Server, SAP HANA, IBM DB2, Hologres, OceanBase (Oracle tenant), Greenplum, Amazon RDS for PostgreSQL, Amazon RDS for SQL Server, Amazon RDS for Oracle, Amazon RDS for DB2, Amazon Redshift, DM (DaMeng), or openGauss, configure the collection range based on the schema, which is the database name under the data source instance. Select All Schemas or Specified Schema.
-
All Schemas: Dynamically retrieves all schemas with query permissions based on the data source configuration.
-
Specified Schema: Specifies other schemas with permissions based on the data source configuration or quickly fills in the default schema with one click. Custom schema names are case-sensitive.
-
-
When the data source is Quick BI, configure the collection range based on workspace. Select all workspaces or specified workspaces.
-
All Workspaces: Dynamically retrieves all workspaces with query permissions based on the application system configuration.
-
Specified Workspace: Specifies other workspaces with permissions based on the application system configuration.
-
Note-
When the collection range is for Hive, StarRocks data sources, the system collects the most recent 100,000 partitions per partitioned table based on creation time.
-
When the data source is OceanBase, the collection range is determined by the tenant mode configured for the data source. MySQL tenant collects metadata based on Database, while Oracle tenant collects metadata based on Schema.
Collection Object Type
Selected by default and cannot be modified. When the data source is a data source, supported types include Tables, Views, and Fields. When the data source is an application system, the supported type is Dashboards.
Note-
When the data source is Elasticsearch, indexes are collected as tables and index aliases are collected as views.
-
When the data source is StarRocks, synchronized materialized views are not supported.
Source System
Only supported when the data source is a data source. Select the source system to which the collected metadata belongs. This is used for asset object filtering, source system lineage display, and other scenarios. For information about how to create a source system, see Create and manage source systems.
Automatic Data Sampling
Available if data sampling is enabled in Administration > Metadata > Sampling Configuration, the trigger scenario includes metadata collection, and data preview is supported. When enabled, sample data is automatically collected during execution based on the collection scope defined in Sampling Configuration > Data Source. You can modify the collection scope.
-
-
Click Next to configure the collection strategy.
Parameter
Description
Data Update Strategy
New/Changed Metadata
Compared with the previous collection, if there is new or updated data in the source system, the system will Add New Metadata And Update Changed Metadata. For dashboards, if a work is modified but not published (status is "Saved but not published"), the system retains the previously collected published data without updating it.
Deleted Metadata
Compared with the previous collection, if there is deleted data in the source system, you can choose Delete from metadata list and asset list or Ignore deletion operation. For dashboards, you can choose If The Work Status Changes From "Published" To "Offline", Treat As Deleted or Ignore Deletion Operation.
-
Delete from metadata list and asset list/If the work status changes from "Published" to "Offline", treat as deleted: Synchronously delete the collected metadata information, which cannot be recovered after deletion.
-
Ignore deletion operation: Ignore the deletion operation in the source system. You can still view the object details and historical versions in the metadata list and asset list, and you can manually delete them later.
Data Collection Schedule
Collection Frequency
Controls the frequency of task collection. Supports Scheduled Collection and Manual Collection.
-
Scheduled Collection: Automatically runs collection at the configured schedule time. Suitable for scenarios with high timeliness requirements. Supports Daily, Weekly, and Monthly schedules. The configurable start time range is 00:00 to 23:59. For Monthly schedules, you can select Last day of month.
When the system time zone (configured in User Center) differs from the scheduling time zone (configured in Management Center > System Settings > Basic Settings), both time zones are displayed. The system automatically converts the scheduled collection time to the scheduling time zone and runs accordingly.
-
Manual Collection: Requires manual triggering. Suitable for scenarios where metadata changes infrequently and resource conservation is desired.
Runtime Configuration
Error Retry
For failed collection instances, determine whether to rerun based on the configured Retry Count and Retry Interval.
-
Retry Count: Whether to automatically retry after a collection instance fails and the maximum number of retries. The default is 1, configurable from 1 to 10.
-
Retry Interval: The interval between retries. The default is 5 minutes, configurable from 1 to 60 minutes.
NoteError retry and scheduled collection may conflict. If the next collection time is reached while the previous task is still running, the next scheduled collection is automatically delayed. You can manually terminate the task in the collection instance list. For more information, see View and manage collection instances.
Runtime Timeout
If the total running time of a collection task (from start to end, excluding resource waiting and scheduling waiting time) exceeds the threshold, the system automatically terminates it and marks it as failed. Configurable from 0 to 24 hours, with up to one decimal place.
Schedule Resource
The collection task uses the resource quota of this resource group when scheduled. To prevent high concurrency from consuming too many resources and affecting other system tasks, all collection tasks across all tenants share a unified concurrency limit. Allocate scheduling resources accordingly. Select resource groups with a status of Normal under the current tenant.
The data source network environment and the scheduling resource group network environment must be interconnected. Otherwise, the collection task cannot run. After selection, click Test Connection to verify network connectivity. If the test fails, click View Log to see the failure reason.
Connection Configuration
View the connection configuration of the selected collection source as a reference for collection frequency and timing. For more information, see Data sources supported by Dataphin.
NoteThe connection configuration applies to offline integration tasks, global quality monitoring rules, and metadata collection tasks.
-
-
Click OK to complete the creation of the collection task.
Manage collection tasks
-
The Collection Task page displays task information, including name, data source and encoding, data source type, collection method, most recent collection status and time, description, owner, effective status, task status, and last update time. Click Datasource Management in the upper right corner to go to the Management Center > Data Source page.
Task Status: The task status determines which operations are available. The following table lists the operations for each status.
Task Status
Operations
Normal
View, Edit, Temporary Manual Execution (supported for scheduled collection tasks), Manual Execution (supported for manual tasks), Clone, Delete, View Metadata, View Collection Instances, Enable or Disable Effective Status.
Creation Failed
Retry, View Execution Log, View, Edit, Delete.
Update Failed/Deletion Failed/Enable Failed/Disable Failed
Retry, View Execution Log, View, Edit, Delete, View Metadata, View Collection Instances.
Enabling/Disabling
View.
Modifying the effective status is not supported when enabling or disabling.
Creating/Updating/Deleting
View.
Abnormal
View, Edit, Delete, View Metadata, View Collection Instances.
-
(Optional) Search for collection tasks by task name or data source name, filter tasks you own or effective tasks, or filter by task status, effective status, owner, data source, or collection method.
-
Perform the following operations in the operation column of the target collection task.
Operation
Description
Retry
Rerun failed collection tasks.
View Execution Log
View the execution logs of failed collection tasks.
View
View the configuration of a collection task.
Edit
You cannot modify the data source type or data source. Other changes do not affect the effective status.
Temporary Manual Execution
Only scheduled collection tasks in Normal status support this operation. If the instance from this execution has not finished when the next scheduled run time is reached, data inconsistency may occur. If the task already has a running instance (scheduled or temporary manual), terminate it first before performing this operation.
Manual Execution
Only manual collection tasks in Normal status support this operation. If the task already has a running instance, terminate it first before performing this operation.
Clone
Quickly copy the configuration of a collection task. You must reconfigure the data source and collection range.
Delete
-
Single Delete: You can click
in the operation column and select Delete to delete the collection task. -
Batch Delete: Select the collection tasks you want to delete, and click the
icon at the bottom to batch delete the collection tasks.
NoteDeleting a task does not affect currently running instances. You can manually terminate them if needed. After deletion, no new collection instances are generated. Choose a deletion strategy: Synchronously delete collected metadata or Only delete the task, retain collected metadata.
-
Synchronously delete collected metadata: Delete the metadata collected from the specified data source through this task from both the metadata list and asset list.
-
Only delete the task, retain collected metadata: Delete only the collection task and retain the collected metadata in the metadata list and asset list. If you later create a new collection task with the same data source, it may overwrite the retained metadata.
View Metadata List
Go to the metadata list page, where the system filters metadata related to the data source configured for this task.
View Collection Instances
Go to the collection instance list page, where the system filters instances related to this task.
Modify Effective Status
-
Modify Single Effective Status: You can click the
switch in the effective status column to enable or disable the effective status. -
Batch Modify Effective Status: Select the collection tasks for which you want to modify the effective status, and click the
icon at the bottom to enable or disable the effective status.
NoteAfter enabling, the collection task runs automatically according to the configured schedule. After disabling, currently running or pending instances are not affected, but subsequent instances are not automatically run. You can still run the task manually.
-
What to do next
-
After the collection task completes, view its execution status in the collection instance list. For more information, see View and manage collection instances.
-
After the collection task runs successfully, view the collected metadata in the metadata list. For more information, see View and manage metadata list.