The TiDB input component reads data from a TiDB data source. To sync TiDB data to another data source, configure the TiDB input component as the read source and then set up the target data source.
Prerequisites
-
You have created a TiDB data source. For instructions, see Create a TiDB Data Source.
-
The account used to configure the TiDB input component must have sync-read permission on the data source. If it does not, request permission. For details, see Request, Renew, or Release Data Source Permissions.
Procedure
-
On the Dataphin homepage, in the top menu bar, click Develop, and then click Data Integration.
-
On the Integration page, select a Project. In Dev-Prod mode, also select an environment.
-
In the left navigation pane, click Batch Pipeline. In the Batch Pipeline list, click the offline pipeline that you want to develop. The pipeline configuration page opens.
-
In the upper-right corner of the page, click Component Library to open the Component Library panel.
-
In the left navigation pane of the Component Library panel, click Input. In the list on the right, locate the TiDB component and drag it onto the canvas.
-
Click the
icon in the TiDB Input component card to open the TiDB Input Configuration dialog box. -
In the TiDB Input Configuration dialog box, configure the parameters.
Parameter
Description
Step Name
The name of the TiDB input component. Dataphin generates a default name that you can change. Naming rules:
-
Use only Chinese characters, letters, underscores (_), and digits.
-
Keep the name under 64 characters.
Datasource
Lists all TiDB data sources in Dataphin, regardless of whether you have sync-read permission. Click the
icon to copy the data source name.-
If you do not have sync-read permission for a data source, click Request next to the data source to request permission. For steps, see Request Data Source Permission.
-
If you do not have any TiDB data sources, click Create Data Source to create one. For steps, see Create a TiDB Data Source.
Source Table Count
Select Single Table or Multiple Tables:
-
Single Table: Syncs data from one source table to one target table.
-
Multiple Tables: Syncs data from multiple source tables to one target table using the union algorithm.
Table Matching Method
Only Generic Rule is supported.
NoteThis setting is available only when you select Multiple Tables for Source Table Count.
Table
Select the source table:
-
If you selected Single Table for Source Table Count, search by keyword or enter the full table name and click Exact Search. After you select a table, the system checks its status automatically. Click the
icon to copy the selected table name. -
If you selected Multiple Tables for Source Table Count, add tables as follows:
-
In the input box, enter a table expression to match tables with the same structure.
The system supports enumerated, regex-like, and mixed formats. For example:
table_[001-100];table_102. -
Click Exact Search. In the Confirm Match Details dialog box, review the list of matched tables.
-
Click Confirm.
-
Shard Key (Optional)
Splits data by the specified column to enable concurrent reads. Any source table column can serve as the shard key. For best performance, use a primary key or indexed column.
ImportantIf you select a date-time type, the system performs brute-force splitting across the full time range and concurrency setting. This method does not guarantee even distribution.
Batch Read Size (Optional)
The number of records to read per batch. Setting a batch size, such as 1024 records, reduces round trips to the source database and lowers network latency.
Input Filter (Optional)
Filter conditions for data extraction:
-
Use a static value to extract matching data. Example:
ds=20210101. -
Use a variable parameter to extract part of the data. Example:
ds=${bizdate}.
Output Fields
Lists all fields from the selected table and filtered results. Available actions:
-
Manage Fields: Remove fields you do not need downstream:
-
Remove One Field: Click the
icon in the Actions column to delete a single field. -
Batch Delete: To delete many fields, click Field Management, select multiple fields in the Field Management dialog box, click the
left-moving icon to move the selected input fields to the unselected input fields, and click OK to complete batch field deletion.
-
-
Batch Add: Click Batch Add to add fields in JSON, TEXT, or DDL format.
NoteAfter you click OK, the new configuration overwrites existing field settings.
-
JSON format example:
// Example: [{ "index": 1, "name": "id", "type": "int(10)", "mapType": "Long", "comment": "comment1" }, { "index": 2, "name": "user_name", "type": "varchar(255)", "mapType": "String", "comment": "comment2" }]Noteindex, name, and type specify the column number of the object, the field name after import, and the field type after import, respectively. For example,
"index": 3, "name": "user_id", "type": "String"indicates that the fourth column in the file is imported as a field named user_id with the type String. -
TEXT format example:
// Example: 1,id,int(10),Long,comment1 2,user_name,varchar(255),Long,comment2-
The row delimiter separates field entries. Default is line feed (\n). Supported delimiters: \n, semicolon (;), and period (.).
-
The column delimiter separates field names and types. Default is comma (,). Supported delimiters:
','. Field type is optional and defaults to','.
-
-
DDL format example:
CREATE TABLE tablename ( user_id serial, username VARCHAR(50), password VARCHAR(50), email VARCHAR (255), created_on TIMESTAMP, );
-
-
Add Output Field: Click + Add Output Field. Enter values for Column, Type, and Comment. Select a Mapping Type. Click the
icon to save the row.
-
-
Click OK to finish configuring the TiDB Input Component.