Migrate PolarDB for MySQL to Elasticsearch

Updated at:

Use Data Transmission Service (DTS) to migrate data from PolarDB for MySQL to Elasticsearch.

Prerequisites

  • You have created a destination Elasticsearch cluster. For more information, see Quick start: Create a cluster and retrieve data.

  • The destination Elasticsearch cluster's storage space must exceed that used by the source PolarDB for MySQL instance.

Limitations

Type

Description

Source database limitations

  • Bandwidth requirements: The server that hosts the source database must have sufficient outbound bandwidth. Otherwise, the data migration speed will be affected.

  • The tables to be migrated must have a primary key or a unique constraint, and the values in the columns must be unique. Otherwise, duplicate data may appear in the destination database.

  • If you migrate data at the table level and need to perform edits such as column mapping, a single data migration task supports a maximum of 1,000 tables. To migrate more tables, we recommend splitting them into multiple tasks or configuring a task to migrate the entire database. Otherwise, a request error may occur after you submit the task.

  • If you perform incremental migration:

    • You must enable binary logging and set the loose_polar_log_bin parameter to on. Otherwise, the precheck reports an error and the data migration task cannot start. For more information about how to enable binary logging and modify parameters, see Enable binary logging and Modify parameters.

      Note

      Enabling binary logging for a PolarDB for MySQL cluster consumes storage space and incurs storage fees.

    • The binary logs of the PolarDB for MySQL cluster must be retained for at least 3 days. We recommend a retention period of 7 days. Otherwise, DTS may fail to obtain the binary logs, which can cause the task to fail. In extreme cases, this can lead to data inconsistency or data loss. Issues caused by a binary log retention period shorter than the DTS requirement are not covered by the DTS Service-Level Agreement (SLA).

      Note

      For more information about how to set the retention period for binary logs of a PolarDB for MySQL cluster, see Modify the retention period.

  • Do not perform DDL operations that change the schema of a database or table during schema migration. Otherwise, the data migration task fails.

Other limitations

  • To add a column to a table that is being migrated, you must first modify the mapping of the corresponding index in the destination Elasticsearch cluster, then run the DDL operation on the source database, and finally pause and restart the data migration task.

  • Do not migrate data to a destination index that contains parent-child relationships or a Join field type mapping. This may cause task exceptions or query failures in the destination database.

  • DTS cannot migrate read-only nodes of a source PolarDB for MySQL cluster.

  • DTS does not support the migration of OSS external tables from the source PolarDB for MySQL instance.

  • DTS does not support migrating the following objects: indexes, partitions, views, procedures, functions, triggers, and foreign keys (FKs).

  • DTS does not support primary/standby switchover scenarios for the database instance during full data migration. In such a scenario, reconfigure the migration task promptly.

  • Full data migration consumes read and write resources on the source and destination databases, which can increase their load. We recommend evaluating the performance of both databases and performing the migration during off-peak hours.

  • Do not use tools such as pt-online-schema-change to perform online DDL operations on migration objects in the source database. Otherwise, the migration fails.

  • For columns of the FLOAT or DOUBLE data type, DTS uses ROUND(COLUMN,PRECISION) to read values. If you do not explicitly define the precision, DTS uses a default migration precision of 38 digits for FLOAT and 308 digits for DOUBLE. Make sure the precision meets your requirements.

  • DTS attempts to resume a failed migration task within seven days. Before you perform a workload switchover to the destination instance, you must stop or release the migration instance. Alternatively, you can run the REVOKE command to revoke the write permissions from the account that DTS uses to access the destination instance. This prevents the task from automatically resuming and overwriting data in the destination instance.

  • The data types supported by PolarDB for MySQL and Elasticsearch clusters are different and do not have a one-to-one mapping. Therefore, during initial schema synchronization, DTS maps the source data types to data types supported by the destination database. For more information, see Data type mappings for initial schema synchronization.

  • Development and test specifications of Elasticsearch instances are not supported.

  • If a task fails, DTS support staff will attempt to restore it within eight hours. During restoration, they may restart the task or adjust its parameters.

    Note

    Only DTS task parameters are modified—not database parameters. Parameters that may be adjusted include those listed in Modify instance parameters.

Other notes

DTS periodically executes the CREATE DATABASE IF NOT EXISTS test command in the source database to advance the Binlog position.

Note

If a synchronization object of DTS is an index alias in the destination Elasticsearch instance, data inconsistency may occur when the actual index that the alias points to changes. For example, when data is written, the index alias points to Index A. If the alias later points to Index B during a subsequent synchronous delete operation, the data on Index A cannot be deleted. As a result, data deleted from the source can still be queried in the destination.

Billing

Migration type

Instance configuration fee

Internet traffic fee

Schema migration and full data migration

Free of charge.

When the Access Method parameter of the destination database is set to Public IP Address, you are charged for Internet traffic. For more information, see Billing overview.

Incremental data migration

Charged. For more information, see Billing overview.

Migration types

  • Schema migration

    DTS migrates the schema definitions of the migration objects from the source database to the destination database.

  • Full migration

    DTS migrates all historical data of the specified migration objects from the source database to the destination database.

  • Incremental migration

    After a full migration is complete, DTS migrates incremental data updates from the source database to the destination database. Incremental migration lets you smoothly migrate data without interrupting your self-managed applications.

SQL operations for incremental migration

Type

SQL statement

DML

INSERT, UPDATE, and DELETE

Note

Using an UPDATE statement to remove a column is not a supported operation.

Permission requirements for database accounts

Database

Permission requirements

Creation and authorization

PolarDB for MySQL cluster

Read permission on the objects to be migrated

Create and manage a database account

Data type mapping

  • A source database and an Elasticsearch instance support different data types that cannot always be mapped directly. During structure initialization, DTS maps data types based on those supported by the target Elasticsearch instance. For more information, see Data type mapping for structure initialization.

    Note

    During the DTS schema migration process, DTS does not set the dynamic parameter in mapping. The behavior of this parameter depends on the settings of your Elasticsearch instance. If your source data is of the JSON type, you must ensure that for a specific key, its corresponding values have the same data type across all rows in a table. Otherwise, DTS may encounter synchronization issues. For more information, see dynamic.

  • The mapping between Elasticsearch and a relational database varies by Elasticsearch version.

    Important

    Starting with Elasticsearch 7.0, an index no longer supports multiple types, and types were completely removed in Elasticsearch 8.0. By default, when you configure a synchronization or migration task, DTS maps a table from a relational database to an index in Elasticsearch. You can change this mapping when you configure the objects to synchronize or migrate.

    Elasticsearch 7.0 and later

    Elasticsearch

    Relational database

    index

    table

    document

    row

    field

    column

    mapping

    schema

    Versions before Elasticsearch 7.0

    Elasticsearch

    Relational database

    index

    database

    type

    table

    document

    row

    field

    column

    mapping

    schema

Procedure

  1. Navigate to the migration task list page for the destination region using one of the following methods.

    From the DTS console

    1. Log on to the Data Transmission Service (DTS) console.

    2. In the navigation pane on the left, click Data Migration.

    3. In the upper-left corner of the page, select the region where the migration instance is located.

    From the DMS console

    Note

    The actual operations may vary based on the mode and layout of the DMS console. For more information, see Simple mode console and Customize the layout and style of the DMS console.

    1. Log on to the Data Management (DMS) console.

    2. In the top menu bar, choose Data + AI > Data Transmission (DTS) > Data Migration.

    3. To the right of Data Migration Tasks, select the region where the migration instance is located.

  2. Click Create Task to navigate to the task configuration page.

  3. Configure the source and destination databases.

    Section

    Parameter

    Description

    N/A

    Task Name

    DTS automatically generates a task name. We recommend that you specify a descriptive name for easy identification. The name does not need to be unique.

    Source Database

    Select Existing Connection

    • To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.

      Note

      In the DMS console, this parameter is named Select a DMS database instance..

    • If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select PolarDB for MySQL.

    Access Method

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region of the source PolarDB for MySQL cluster.

    Cross-account

    This example migrates data within the same Alibaba Cloud account. Select No.

    PolarDB Cluster ID

    Select the cluster ID of the source PolarDB for MySQL cluster.

    Database Account

    Enter the database account of the source PolarDB for MySQL cluster. For information about the required permissions, see Permissions required for the database account.

    Database Password

    Enter the password for the database account.

    Encryption

    Select a connection type as needed. For more information about the SSL encryption feature, see Enable SSL encryption.

    Destination Database

    Select Existing Connection

    • To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.

      Note

      In the DMS console, this parameter is named Select a DMS database instance..

    • If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select Elasticsearch.

    Access Method

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region of the destination Elasticsearch cluster.

    Type

    Select Cluster or Serverless based on your requirements.

    Instance ID

    Select the ID of the destination Elasticsearch cluster.

    Database Account

    Enter the database account for the Elasticsearch cluster. The default account is elastic.

    Database Password

    Enter the logon password that you set when you created the Elasticsearch cluster.

    Encryption

    Select HTTP or HTTPS as needed.

  4. After you complete the configuration, click Test Connectivity and Proceed at the bottom of the page.

    Note
    • Ensure that the IP address segment of the DTS service is automatically or manually added to the security settings of the source and destination databases to allow access from DTS servers. For more information, see Add DTS server IP addresses to a whitelist.

    • If the source or destination database is a self-managed database (the Access Method is not Alibaba Cloud Instance), you must also click Test Connectivity in the CIDR Blocks of DTS Servers dialog box that appears.

  5. Configure the task objects.

    1. On the Configure Objects page, configure the objects that you want to migrate.

      Parameter

      Description

      Migration Types

      • If you only need to perform a full migration, select both Schema Migration and Full Data Migration.

      • To perform a migration with no downtime, select Schema Migration, Full Data Migration, and Incremental Data Migration.

      Note
      • If you do not select Schema Migration, you must ensure that a database and tables to receive the data exist in the destination database. You can also use the object name mapping feature in the Selected Objects box as needed.

      • If you do not select Incremental Data Migration, do not write new data to the source instance during data migration to ensure data consistency.

      Processing Mode for Existing Destination Tables

      • Precheck and Report Errors: DTS checks for tables in the destination database with the same names as those in the source. The precheck passes if no such tables exist. Otherwise, the precheck fails, and the data migration task does not start.

        Note

        If you cannot delete or rename the tables with the same names in the destination database, you can change their names in the destination database. For more information, see object name mapping.

      • Ignore Errors and Proceed: Skips the precheck for tables with the same names.

        Warning

        Selecting Ignore Errors and Proceed may cause data inconsistency and risk service disruption. For example:

        • If the table schemas are consistent and a record in the destination database has the same primary key value as a record in the source database:

          • During full data migration, DTS retains the record in the destination cluster. The record from the source database is not migrated.

          • During incremental data migration, DTS does not retain the record in the destination cluster. The record from the source database overwrites the one in the destination database.

        • If the table schemas are inconsistent, data initialization may fail, only partial data may be migrated, or the migration may fail.

      Index Name

      • Table Name

        If you select Table Name, the index created in the destination Elasticsearch instance has the same name as the source table.

      • Database Name_Table Name

        If you select Database Name_Table Name, the index created in the destination Elasticsearch instance is named in the DatabaseName_TableName format.

      Case Policy for Destination Object Names

      Configure case sensitivity for database, table, and column names for the migrated objects in the destination instance. By default, DTS default policy is selected. You can also choose to maintain the same case as the source or destination database. For more information, see Case sensitivity of object names in the destination database.

      Source Objects

      Select one or more objects from the Source Objects section. Click the Rightwards arrow icon and add the objects to the Selected Objects section.

      Note

      You can select databases or tables as migration objects. If you select tables, other objects such as views, triggers, and stored procedures are not migrated to the destination database.

      Selected Objects

      Note
      • Only underscores (_) are supported as special characters in index names and type names.

      • To filter data with WHERE conditions or specify post-migration details like index, type, and column names, right-click a table in the Selected Objects box. Configure the settings in the dialog box that appears. For more information, see Configure filter conditions.

      • To select SQL operations to migrate at the database or table level, right-click a migration object in the Selected Objects box and select the desired SQL operations in the dialog box that appears. For information about supported operations, see SQL operations supported for incremental data migration.

    2. Click Next: Advanced Settings to configure advanced parameters.

      Parameter

      Description

      Dedicated Cluster for Task Scheduling

      By default, DTS schedules tasks on a shared cluster. You do not need to select one. If you want more stable tasks, you can purchase a dedicated cluster to run DTS migration tasks.

      Retry Time for Failed Connections

      After the migration task starts, if the connection to the source or destination database fails, DTS reports an error and immediately begins to retry the connection. The default retry duration is 720 minutes. You can customize the retry time to a value from 10 to 1440 minutes. We recommend that you set the duration to more than 30 minutes. If DTS reconnects to the source and destination databases within the specified duration, the migration task automatically resumes. Otherwise, the task fails.

      Note
      • For multiple DTS instances that share the same source or destination, the network retry time is determined by the setting of the last created task.

      • Because you are charged for the task during the connection retry period, we recommend that you customize the retry time based on your business needs, or release the DTS instance as soon as possible after the source and destination database instances are released.

      Retry Time for Other Issues

      After the migration task starts, if a non-connectivity issue, such as a DDL or DML execution exception, occurs in the source or destination database, DTS reports an error and immediately begins to retry the operation. The default retry duration is 10 minutes. You can customize the retry time to a value from 1 to 1440 minutes. We recommend that you set the duration to more than 10 minutes. If the related operations succeed within the specified retry duration, the migration task automatically resumes. Otherwise, the task fails.

      Important

      The value of Retry Time for Other Issues must be less than the value of Retry Time for Failed Connections.

      Enable Throttling for Full Data Migration

      During full migration, DTS consumes read and write resources on the source and destination databases, which may increase the database load. If required, you can enable throttling for the full migration task. You can set Queries per second (QPS) to the source database, RPS of Full Data Migration, and Data migration speed for full migration (MB/s) to reduce the load on the destination database.

      Note
      • This configuration item is available only if you select Full Data Migration for Migration Types.

      • You can also adjust the full migration speed after the migration instance is running.

      Enable Throttling for Incremental Data Migration

      If required, you can also choose to set speed limits for the incremental migration task. You can set RPS of Incremental Data Migration and Data migration speed for incremental migration (MB/s) to reduce the load on the destination database.

      Note
      • This configuration item is available only if you select Incremental Data Migration for Migration Types.

      • You can also adjust the incremental migration speed after the migration instance is running.

      Environment Tag

      You can select an environment tag to identify the instance based on your requirements. This parameter is not required for this example.

      Shard Configuration

      Set the number of primary and replica shards for the index based on the maximum shard configuration for an index in the destination Elasticsearch cluster.

      String Index

      Specifies how to index strings migrated to the destination Elasticsearch cluster.

      • analyzed: Analyzes the strings before indexing. You must also select a specific analyzer. For more information about analyzer types and their functions, see Analyzers.

      • not analyzed: Indexes the strings directly by using their original values without analysis.

      • no: Does not index the strings.

      Time Zone

      Select the time zone for time-related data types (such as DATETIME and TIMESTAMP) migrated to the destination Elasticsearch instance.

      Note

      If the time-related data types in the destination instance do not require a time zone, you must set the document type (type) for these data types in the destination instance before synchronization.

      DOCID

      By default, the DOCID is the primary key of the table. If the table has no primary key, the DOCID is an ID column automatically generated by Elasticsearch.

      Whether to delete SQL operations on heartbeat tables of forward and reverse tasks

      Choose whether DTS writes heartbeat SQL information to the source database while the instance is running.

      • Yes: Does not write heartbeat SQL information to the source database. The DTS instance may display latency.

      • No: Writes heartbeat SQL information to the source database. This may interfere with source database operations like physical backups and cloning.

      Configure ETL

      Based on your business needs, select whether to configure the ETL feature to process data.

      • Yes: Configures the ETL feature. You must also enter data processing statements in the text box.

      • No: Does not configure the ETL feature.

      Monitoring and Alerting

      Select whether to set alerts and receive alert notifications based on your business needs.

      • No: Does not set an alert.

      • Yes: Configure alerts by setting an alert threshold and an alert contact. If a migration fails or the latency exceeds the threshold, the system sends an alert notification.

    3. After completing the preceding settings, click Next: Configure Database and Table Fields to set the _routing policy and _id value for the tables to be migrated to the destination Elasticsearch instance.

      Parameter

      Description

      Set _routing

      The _routing parameter routes a document to a specific shard in the destination Elasticsearch instance. For more information, see _routing.

      • Select Yes to specify custom columns for routing.

      • Select No to use the _id value for routing.

      Note

      If the destination Elasticsearch instance is version 7.x, select No.

      Value of _id

      • Primary key column

        A composite primary key is merged into a single column.

      • Business key

        If you select Business key, you must also specify the business key column.

  6. Save the task and run a precheck.

    • To view the parameters for configuring this instance when you call the API operation, move the pointer over the Next: Save Task Settings and Precheck button and click Preview OpenAPI parameters in the bubble that appears.

    • If you do not need to view or have finished viewing the API parameters, click Next: Save Task Settings and Precheck at the bottom of the page.

    Note
    • Before the migration task starts, DTS performs a precheck. The task starts only after it passes the precheck.

    • If the precheck fails, click View Details next to the failed check item, fix the issue based on the prompt, and then run the precheck again.

    • If a warning is reported during the precheck:

      • For check items that cannot be ignored, click View Details next to the failed item, fix the issue based on the prompt, and then run the precheck again.

      • For check items that can be ignored, you can click Confirm Alert Details, Ignore, OK, and Precheck Again to skip the alert item and run the precheck again. If you choose to ignore a warning, it may cause issues such as data inconsistency and pose risks to your business.

  7. Purchase the instance.

    1. When the Success Rate is 100%, click Next: Purchase Instance.

    2. On the Purchase page, select the link specification for the data migration instance. For more information, see the following table.

      Category

      Parameter

      Description

      New Instance Class

      Resource Group Settings

      Select the resource group to which the instance belongs. The default value is default resource group. For more information, see What is Resource Management?

      Instance Class

      DTS provides migration specifications with different performance levels. The link specification affects the migration speed. You can select a specification based on your business scenario. For more information, see Data migration link specifications.

    3. After the configuration is complete, read and select Data Transmission Service (Pay-as-you-go) Service Terms.

    4. Click Buy and Start. In the OK dialog box that appears, click OK.

      You can view the progress of the migration task on the Data Migration Tasks list page.

      Note
      • If the migration task does not include incremental migration, it stops automatically after the full migration is complete. After the task stops, its Status changes to Completed.

      • If the migration task includes incremental migration, it does not stop automatically. The incremental migration task continues to run. While the incremental migration task is running, the Status of the task is Running.