All Products
Search
Document Center

Data Transmission Service:Synchronize PolarDB-X 1.0 to Elasticsearch

Last Updated:Aug 18, 2026

This topic describes how to use DTS to synchronize data from a PolarDB-X 1.0 instance to an Elasticsearch cluster.

Prerequisites

  • The source PolarDB-X 1.0 instance must use ApsaraDB RDS for MySQL as its storage type. This includes custom and separately purchased ApsaraDB RDS instances. PolarDB for MySQL is not supported.

  • You have created a destination Elasticsearch cluster. For more information, see Create an Alibaba Cloud Elasticsearch instance.

  • The available storage space of the destination Elasticsearch cluster must exceed the total data size of the source PolarDB-X 1.0 instance.

Considerations

Type

Description

Source database limits

  • Tables to be synchronized must have a PRIMARY KEY or UNIQUE constraint. Schema synchronization is not supported for tables with only a UNIQUE constraint — use a PRIMARY KEY constraint instead. The key fields must be unique. Otherwise, duplicate data may appear in the destination database.

  • If you synchronize data at the table level and need to edit objects such as table or column name mappings, a single synchronization task supports up to 1,000 tables. If you exceed this limit, an error is reported when you submit the task. To work around this, split the tables into multiple tasks or synchronize the entire database.

  • Binary logs of the RDS for MySQL instances attached to the PolarDB-X 1.0 instance:

    • Binary logging is enabled by default. For more information, see View instance parameters to confirm that binlog_row_image is set to full. Otherwise, an error is reported during the precheck, and the data synchronization task cannot start.

    • For incremental synchronization, DTS requires that local binary logs of the source database are retained for more than 24 hours. For full-plus-incremental synchronization, local binary logs must be retained for at least 7 days. You can change the retention period to more than 24 hours after full synchronization is complete. If DTS cannot obtain the binary logs, the task may fail and, in extreme cases, cause data inconsistency or loss. Issues caused by insufficient binary log retention are not covered by the DTS Service-Level Agreement (SLA).

  • Source database operation limitations:

    • To switch the network type of the PolarDB-X 1.0 instance during synchronization, update the network connection information of the synchronization link after the switch is successful.

    • During synchronization, do not perform operations on the source instance such as scaling (for example, scaling its attached ApsaraDB RDS for MySQL instance, or changing the distribution of physical tables that correspond to logical tables in ApsaraDB RDS for MySQL even if the ApsaraDB RDS for MySQL instance is not scaled), migrating hot tables, changing the shard key, or making DDL changes. Otherwise, the data synchronization task will fail or data inconsistencies will occur.

    • If you change the partition key of a source table, and the primary key of the corresponding destination table does not include the partition key, data loss may occur in the destination database.

      • Reason: A PolarDB-X 1.0 instance has multiple ApsaraDB RDS for MySQL instances attached to it. DTS maintains independent synchronization links for each ApsaraDB RDS for MySQL instance. Changing a partition key causes data to be deleted from one ApsaraDB RDS for MySQL instance and inserted into another. Because these links execute independently, the order of the DELETE and INSERT operations can be disrupted. If the primary key of the destination database does not include the partition key, a DELETE operation that is executed later can incorrectly delete the newly inserted data.

      • Recommendation: To prevent data loss from out-of-order operations, ensure the primary key of the destination table includes the partition key columns of the source table.

    • If you perform only full data synchronization, do not write new data to the source instance. Otherwise, data inconsistency occurs between the source and destination. To maintain real-time data consistency, select schema synchronization, full data synchronization, and incremental data synchronization.

    • Do not run DDL operations that change database or table schemas during schema synchronization or full synchronization. Otherwise, the synchronization task fails.

      Note

      During full synchronization, DTS queries the source database. This creates metadata locks that may block DDL operations on the source database.

  • The source PolarDB-X 1.0 instance must be version 5.2 or later.

Other limits

  • DTS does not synchronize objects of the following types: index, partition, view, procedure, function, trigger, and foreign key.

  • To add a column to a synchronized table, you must first modify the table mapping in the Elasticsearch instance. Then, run the corresponding DDL operation on the source database. Finally, pause and start the data synchronization task.

  • You cannot synchronize data to Elasticsearch indexes that contain parent-child relationships or Join field types. Doing so may cause task errors or query failures in the destination.

  • Incremental data synchronization relies on continuous XA transactions in the source PolarDB-X 1.0 instance to ensure data consistency. If this continuity is disrupted, for example when modifying synchronization objects or during disaster recovery of the incremental data capture module, uncommitted XA transactions may be lost.

  • Before you synchronize data, evaluate the performance of the source and destination databases. Synchronize data during off-peak hours. Otherwise, during full data synchronization, DTS consumes some read and write resources from both the source and destination databases, which may increase the database load.

  • DTS attempts to resume data synchronization tasks that failed within the last seven days. Therefore, before you switch your workload to the destination instance, you must stop or release the failed tasks. Alternatively, run the REVOKE command to revoke the write permissions from the account that DTS uses to access the destination instance. This prevents an automatically resumed task from overwriting data in the destination instance with data from the source.

  • Development and test specifications of Elasticsearch instances are not supported.

  • If a task fails, DTS support staff will attempt to restore it within eight hours. During restoration, they may restart the task or adjust its parameters.

    Note

    Only DTS task parameters are modified—not database parameters. Parameters that may be adjusted include those listed in Modify instance parameters.

Other notes

DTS periodically updates the dts_health_check.ha_health_check table in the source database to advance the binary log position.

Note

If the objects to be synchronized by DTS are index aliases in the destination Elasticsearch, data inconsistency may occur when the actual index that an index alias points to changes. For example, the index alias points to index A when data is written. However, during subsequent synchronization of delete operations, the index alias points to index B. In this case, the data on index A cannot be deleted, and the data that has been deleted from the source can still be queried in the destination.

Billing

Synchronization type

Pricing

Schema synchronization and full data synchronization

Free of charge.

Incremental data synchronization

Charged. For more information, see Billing overview.

SQL operations for incremental synchronization

Operation type

SQL statement

DML

INSERT, UPDATE, and DELETE

Note

Column removal using UPDATE is not synchronized.

Permissions for database accounts

Instance

Required permissions

References

Source PolarDB-X 1.0 instance

Read permission on the objects to be synchronized.

Account management

Destination Elasticsearch instance

The database account must have read and write permissions. Typically, this is the elastic account.

Data type mapping

  • A source database and an Elasticsearch instance support different data types that cannot always be mapped directly. During structure initialization, DTS maps data types based on those supported by the target Elasticsearch instance. For more information, see Data type mapping for structure initialization.

    Note

    During the DTS schema migration process, DTS does not set the dynamic parameter in mapping. The behavior of this parameter depends on the settings of your Elasticsearch instance. If your source data is of the JSON type, you must ensure that for a specific key, its corresponding values have the same data type across all rows in a table. Otherwise, DTS may encounter synchronization issues. For more information, see dynamic.

  • The mapping between Elasticsearch and a relational database varies by Elasticsearch version.

    Important

    Starting with Elasticsearch 7.0, an index no longer supports multiple types, and types were completely removed in Elasticsearch 8.0. By default, when you configure a synchronization or migration task, DTS maps a table from a relational database to an index in Elasticsearch. You can change this mapping when you configure the objects to synchronize or migrate.

    Elasticsearch 7.0 and later

    Elasticsearch

    Relational database

    index

    table

    document

    row

    field

    column

    mapping

    schema

    Versions before Elasticsearch 7.0

    Elasticsearch

    Relational database

    index

    database

    type

    table

    document

    row

    field

    column

    mapping

    schema

Precautions

Procedure

  1. Go to the data synchronization task list page in the destination region. You can do this in one of two ways.

    DTS console

    1. Log on to the DTS console.

    2. In the navigation pane on the left, click Data Synchronization.

    3. In the upper-left corner of the page, select the region where the synchronization instance is located.

    DMS console

    Note

    The actual steps may vary depending on the mode and layout of the DMS console. For more information, see Simple mode console and Customize DMS console layout and style.

    1. Log on to the DMS console.

    2. In the top menu bar, choose Data + AI > DTS (DTS) > Data Synchronization.

    3. To the right of Data Synchronization Tasks, select the region of the synchronization instance.

  2. Click Create Task to open the task configuration page.

  3. Configure the source and destination databases.

    Category

    Parameter

    Description

    N/A

    Task Name

    DTS automatically generates a task name. We recommend that you specify a descriptive name for easy identification. The name does not need to be unique.

    Source Database

    Select Existing Connection

    • Select the registered database instance with DTS from the drop-down list. The database information below is automatically configured.

      Note

      In the DMS console, this configuration item is Select a DMS database instance.

    • If you have not registered the database instance or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select PolarDB-X 1.0.

    Access Method

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region where the source PolarDB-X 1.0 instance is located.

    Replicate Data Across Alibaba Cloud Accounts

    This example synchronizes data within the same Alibaba Cloud account. Select No.

    Instance ID

    Select the instance ID of the source PolarDB-X 1.0 instance.

    Database Account

    Enter the database account of the source PolarDB-X 1.0 instance.

    Database Password

    Enter the password for the specified database account.

    Destination Database

    Select Existing Connection

    • Select the registered database instance with DTS from the drop-down list. The database information below is automatically configured.

      Note

      In the DMS console, this configuration item is Select a DMS database instance.

    • If you have not registered the database instance or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select Elasticsearch.

    Access Method

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region where the destination Elasticsearch instance is located.

    Type

    Select Cluster or Serverless based on your business requirements.

    Instance ID

    Select the instance ID of the destination Elasticsearch instance.

    Database Account

    Enter the account used to connect to the destination Elasticsearch instance. This is the Login name that you specified when you created the Elasticsearch instance. The default account is elastic.

    Database Password

    Enter the password for the specified database account.

    Encryption

    Select HTTP or HTTPS as needed.

  4. After completing the configuration, click Test Connectivity and Proceed at the bottom of the page.

    Note
    • Ensure that you add the CIDR blocks of the DTS servers (either automatically or manually) to the security settings of both the source and destination databases to allow access. For more information, see Add the IP address whitelist of DTS servers.

    • If the source or destination is a self-managed database (i.e., the Access Method is not Alibaba Cloud Instance), you must also click Test Connectivity in the CIDR Blocks of DTS Servers dialog box.

  5. Configure the task objects.

    1. On the Configure Objects page, specify the objects to synchronize.

      Parameter

      Description

      Synchronization Types

      DTS always selects Incremental Data Synchronization. By default, you must also select Schema Synchronization and Full Data Synchronization. After the precheck, DTS initializes the destination cluster with the full data of the selected source objects, which serves as the baseline for subsequent incremental synchronization.

      Processing Mode of Conflicting Tables

      • Precheck and Report Errors: Checks whether the destination database contains an index with the same name. If no index with the same name exists, the precheck passes. If an index with the same name exists, DTS reports an error during the precheck, and the synchronization task does not start.

        Note

        If you cannot delete or rename the conflicting index in the destination database, you can specify a different name for the synchronized object in the destination instance to avoid name conflicts.

      • Ignore Errors and Proceed: Skips the check for indexes with the same name in the destination database.

        Warning

        If you select Ignore Errors and Proceed, data inconsistency may occur, which poses risks to your business. For example:

        • If the mapping structures are the same and a record in the destination database has the same primary key value as a record in the source database, DTS retains the destination record during initialization but overwrites it during incremental data synchronization.

        • If the mapping structures are different, data initialization may fail, only partial columns may be synchronized, or the synchronization may fail entirely.

      Index Name

      • If you select Table Name, the index in the destination Elasticsearch instance is named after the source table.

      • If you select Database Name_Table Name, the destination Elasticsearch index is named in the database_name_table_name format.

      Note

      This index name mapping configuration applies to all tables.

      Capitalization of Object Names in Destination Instance

      Configure the case-sensitivity policy for database, table, and column names in the destination instance. By default, the DTS default policy is selected. You can also choose to use the default policy of the source or destination database. For more information, see Case policy for destination object names.

      Source Objects

      In the Source Objects section, click the objects that you want to synchronize, and then click 向右小箭头 to add them to the Selected Objects section.

      Note

      We recommend that you select objects at the table level. If you select an entire database as a synchronization object, changes to add or delete tables in the source database are not synchronized to the destination database.

      Selected Objects

      • To rename a single object in the destination instance, right-click the object in the Selected Objects box. For more information, see Map a single object name.

      • To rename multiple objects in bulk, click Batch Edit in the upper-right corner of the Selected Objects box. For more information, see Map multiple object names in bulk.

      Note
      • Only underscores (_) are supported as special characters in index and type names.

      • To specify WHERE conditions to filter data, right-click a table in the Selected Objects section and specify the conditions in the dialog box that appears. For more information, see Set filter conditions.

      • If you use the object name mapping feature, synchronization of dependent objects might fail.

    2. Click Next: Advanced Settings.

      Parameter

      Description

      Dedicated Cluster for Task Scheduling

      By default, DTS uses a shared cluster for tasks, so you do not need to make a selection. For greater task stability, you can purchase a dedicated cluster to run the DTS synchronization task. For more information, see What is a DTS dedicated cluster?.

      Retry Time for Failed Connections

      If the connection to the source or destination database fails after the synchronization task starts, DTS reports an error and immediately begins to retry the connection. The default retry duration is 720 minutes. You can customize the retry time to a value from 10 to 1,440 minutes. We recommend a duration of 30 minutes or more. If the connection is restored within this period, the task resumes automatically. Otherwise, the task fails.

      Note
      • If multiple DTS instances (e.g., Instance A and B) share a source or destination, DTS uses the shortest configured retry duration (e.g., 30 minutes for A, 60 for B, so 30 minutes is used) for all instances.

      • DTS charges for task runtime during connection retries. Set a custom duration based on your business needs, or release the DTS instance promptly after you release the source/destination instances.

      Retry Time for Other Issues

      If a non-connection issue (e.g., a DDL or DML execution error) occurs, DTS reports an error and immediately retries the operation. The default retry duration is 10 minutes. You can also customize the retry time to a value from 1 to 1,440 minutes. We recommend a duration of 10 minutes or more. If the related operations succeed within the set retry time, the synchronization task automatically resumes. Otherwise, the task fails.

      Important

      The value of Retry Time for Other Issues must be less than that of Retry Time for Failed Connections.

      Enable Throttling for Full Data Synchronization

      During full data synchronization, DTS consumes read and write resources from the source and destination databases, which can increase their load. To mitigate pressure on the destination database, you can limit the migration rate by setting Queries per second (QPS) to the source database, RPS of Full Data Migration, and Data migration speed for full migration (MB/s).

      Note

      Enable Throttling for Incremental Data Synchronization

      You can also limit the incremental synchronization rate to reduce pressure on the destination database by setting RPS of Incremental Data Synchronization and Data synchronization speed for incremental synchronization (MB/s).

      Environment Tag

      Select an environment tag to identify the instance based on your requirements.

      Shard Configuration

      Set the number of primary shards and replica shards for the index based on the maximum shard configuration of indexes in the destination Elasticsearch instance.

      String Index

      Specifies how to index strings synchronized to the destination Elasticsearch instance.

      • analyzed: The string is analyzed before being indexed. You must also select a specific analyzer. For more information about analyzer types and their functions, see Analyzers.

      • not_analyzed: The string is indexed with its original value without analysis.

      • no: The string is not indexed.

      Time Zone

      When DTS synchronizes time-related data types such as DATETIME and TIMESTAMP to the destination Elasticsearch instance, you can select the time zone to be included.

      Note

      If time-related data types in the destination instance do not require a time zone, you must set the document type (type) for these data types in the destination instance beforehand.

      DOCID

      The DOCID defaults to the primary key of the table. If the table has no primary key, the DOCID is an ID column that the Elasticsearch instance automatically generates.

      Configure ETL

      Choose whether to enable the extract, transform, and load (ETL) feature. For more information, see What is ETL? Valid values:

      Monitoring and Alerting

      Choose whether to set up alerts. If the synchronization fails or the latency exceeds the specified threshold, DTS sends a notification to the alert contacts.

    3. After you complete the preceding configurations, click Next: Configure Database and Table Fields at the bottom of the page. Then, set the _routing policy and the _id value for the tables that you want to synchronize to the destination Elasticsearch instance.

      Type

      Description

      Set _routing

      You can use the _routing feature to store a document on a specific shard of the destination Elasticsearch instance. For more information, see _routing.

      • If you select Yes, you can specify a custom column for routing.

      • If you select No, the value of _id is used for routing.

      Note

      If the destination Elasticsearch instance runs version 7.x, you must select No.

      _routing Column

      Select the column to use for routing.

      Note

      This parameter is required only if you set Set _routing to Yes.

      Value of _id

      Select the column to use as the document ID.

  6. Save the task and perform a precheck.

    • To view the parameters for configuring this instance via an API operation, hover over the Next: Save Task Settings and Precheck button and click Preview OpenAPI parameters in the tooltip.

    • If you have finished viewing the API parameters, click Next: Save Task Settings and Precheck at the bottom of the page.

    Note
    • Before a synchronization task starts, DTS performs a precheck. You can start the task only if the precheck passes.

    • If the precheck fails, click View Details next to the failed item, fix the issue as prompted, and then rerun the precheck.

    • If the precheck generates warnings:

      • For non-ignorable warning, click View Details next to the item, fix the issue as prompted, and run the precheck again.

      • For ignorable warnings, you can bypass them by clicking Confirm Alert Details, then Ignore, and then OK. Finally, click Precheck Again to skip the warning and run the precheck again. Ignoring precheck warnings may lead to data inconsistencies and other business risks. Proceed with caution.

  7. Purchase the instance.

    1. When the Success Rate reaches 100%, click Next: Purchase Instance.

    2. On the Purchase page, select the billing method and link specifications for the data synchronization instance. For more information, see the following table.

      Category

      Parameter

      Description

      New Instance Class

      Billing Method

      • Subscription: You pay upfront for a specific duration. This is cost-effective for long-term, continuous tasks.

      • Pay-as-you-go: You are billed hourly for actual usage. This is ideal for short-term or test tasks, as you can release the instance at any time to save costs.

      Resource Group Settings

      The resource group to which the instance belongs. The default is default resource group. For more information, see What is Resource Management?.

      Instance Class

      DTS offers synchronization specifications at different performance levels that affect the synchronization rate. Select a specification based on your business requirements. For more information, see Data synchronization link specifications.

      Subscription Duration

      In subscription mode, select the duration and quantity of the instance. Monthly options range from 1 to 9 months. Yearly options include 1, 2, 3, or 5 years.

      Note

      This option appears only when the billing method is Subscription.

    3. Read and select the checkbox for Data Transmission Service (Pay-as-you-go) Service Terms.

    4. Click Buy and Start, and then click OK in the OK dialog box.

      You can monitor the task progress on the data synchronization page.