Migrate from PolarDB-X 2.0 to Elasticsearch

Updated at:

You can use DTS to migrate data from a PolarDB-X instance to an Elasticsearch instance.

Prerequisites

  • A source PolarDB-X 2.0 instance has been created.

  • A destination Elasticsearch instance has been created. For more information, see Create an Alibaba Cloud Elasticsearch instance.

  • For the supported versions of the source and destination instances, see Overview of migration solutions.

  • The destination Elasticsearch instance must have more storage space than the source PolarDB-X 2.0 instance.

Limitations

Note
  • DTS migrates foreign keys during schema migration.

  • During full and incremental data migration, DTS temporarily disables constraint checks and foreign key cascades at the session level. Cascade updates or deletes on the source database while the task runs may cause data inconsistency.

Type

Description

Source database limitations

  • The server that hosts the source database must have sufficient outbound bandwidth. Insufficient bandwidth slows down data migration.

  • Enterprise Edition PolarDB-X 2.0 read-only instances are not supported as a source database.

  • Tables to be migrated must have a PRIMARY KEY or a UNIQUE constraint, and the fields in the key or constraint must be unique. Otherwise, this may cause duplicate data in the destination database.

  • If you migrate objects at the table level and need to edit them, for example, by mapping column names, a single data migration task supports a maximum of 1,000 tables. If you exceed this limit, the task submission will fail. In this case, split the tables into multiple data migration tasks or configure a task to migrate the entire database.

  • For incremental data migration, the source database must meet the following binary log requirements:

    • Enable the binary log feature, and set the binlog_row_image parameter to full. Otherwise, the precheck fails and the data migration task cannot start.

    • For incremental data migration tasks, DTS requires that the binary logs of the source database be retained for at least 24 hours. For tasks that include both full data migration and incremental data migration, the binary logs must be retained for at least 7 days. You can change the retention period to 24 hours after the full data migration is complete. If the binary logs are not retained for the required period, DTS may fail to obtain them, which can cause task failures or even data inconsistency and loss. The DTS SLA does not cover issues caused by an insufficient binary log retention period.

  • If a table name in the PolarDB-X 2.0 instance contains uppercase letters, only schema migration is supported for that table.

  • Operational limitations on the source database:

    • During schema migration and full data migration, do not perform DDL operations that change the database or table schema. Otherwise, the data migration task will fail.

      Note

      During full data migration, DTS queries the source database. This action places a metadata lock, which may block DDL operations on the source database.

    • If you need to change the network type of the PolarDB-X 2.0 instance during migration, you must update the network connection information of the migration task after the change is complete.

    • If you perform only full data migration, do not write new data to the source database during the migration. Otherwise, data will become inconsistent between the source and destination databases. To maintain real-time data consistency, select schema migration, full data migration, and incremental data migration.

  • Migration of table groups (TABLEGROUP) and databases or tables with the Locality attribute is not supported.

  • Migration of tables whose names are reserved words, such as select, is not supported.

  • In a PolarDB-X 2.0 instance, database partitions in DRDS mode are not supported for synchronization.

  • During the operation of a DTS migration task, changing the type of a broadcast table in the source PolarDB-X 2.0 instance (for example, changing a broadcast table to a regular table or a sharded table) is not supported. To change the table type, stop the migration task first, and then reconfigure the migration task after the change is complete.

Other limits

  • If you need to add columns to a table during migration, you must first update the corresponding index mapping in your Elasticsearch instance. Then, run the DDL operation on the source database. Finally, pause and restart the data migration task.

  • Do not migrate data to a destination index that contains parent-child relationships or a Join field type mapping. This may cause task exceptions or query failures in the destination database.

  • Development and test specifications of Elasticsearch instances are not supported.

  • Evaluate the performance of the source and destination databases before you start migration. We recommend that you perform data migration during off-peak hours because DTS consumes read and write resources on both databases during full data migration, which increases database load.

  • During full data migration, concurrent INSERT operations can cause table fragmentation in the destination database. As a result, tables in the destination database may occupy more storage space than those in the source database.

  • DTS attempts to recover failed tasks for up to seven days. Before you switch your workloads to the destination database, end or release the task, or revoke the write permissions of the DTS account on the destination database by using the REVOKE command. This prevents DTS from overwriting data in the destination database if the task is automatically recovered.

  • If a task fails, DTS support staff will attempt to restore it within eight hours. During restoration, they may restart the task or adjust its parameters.

    Note

    Only DTS task parameters are modified—not database parameters. Parameters that may be adjusted include those listed in Modify instance parameters.

Other precautions

DTS periodically updates the dts_health_check.ha_health_check table in the source database to advance the binlog position.

Note

If a synchronization object of DTS is an index alias in the destination Elasticsearch instance, data inconsistency may occur when the actual index that the alias points to changes. For example, when data is written, the index alias points to Index A. If the alias later points to Index B during a subsequent synchronous delete operation, the data on Index A cannot be deleted. As a result, data deleted from the source can still be queried in the destination.

Billing

Migration type

Instance configuration fee

Internet traffic fee

Schema migration and full data migration

Free of charge.

When the Access Method parameter of the destination database is set to Public IP Address, you are charged for Internet traffic. For more information, see Billing overview.

Incremental data migration

Charged. For more information, see Billing overview.

Migration types

  • Schema migration

    DTS migrates the schema definitions of the migration objects from the source database to the destination database.

  • Full migration

    DTS migrates all historical data of the specified migration objects from the source database to the destination database.

  • Incremental migration

    After a full migration is complete, DTS migrates incremental data updates from the source database to the destination database. Incremental migration lets you smoothly migrate data without interrupting your self-managed applications.

SQL operations for incremental migration

Type

SQL statement

DML

INSERT, UPDATE, and DELETE

Note

Using an UPDATE statement to remove a column is not a supported operation.

Permission requirements for database accounts

Database

Schema migration

Full migration

Incremental migration

Source PolarDB-X 2.0 instance

SELECT permission

SELECT permission

REPLICATION SLAVE, REPLICATION CLIENT, and SELECT permissions on the objects to be migrated.

Note

For information on how to grant permissions, see Account permission issues during data synchronization.

Destination Elasticsearch instance

The database account must have read and write permissions. Typically, this is the elastic account.

Data type mapping

  • A source database and an Elasticsearch instance support different data types that cannot always be mapped directly. During structure initialization, DTS maps data types based on those supported by the target Elasticsearch instance. For more information, see Data type mapping for structure initialization.

    Note

    During the DTS schema migration process, DTS does not set the dynamic parameter in mapping. The behavior of this parameter depends on the settings of your Elasticsearch instance. If your source data is of the JSON type, you must ensure that for a specific key, its corresponding values have the same data type across all rows in a table. Otherwise, DTS may encounter synchronization issues. For more information, see dynamic.

  • The mapping between Elasticsearch and a relational database varies by Elasticsearch version.

    Important

    Starting with Elasticsearch 7.0, an index no longer supports multiple types, and types were completely removed in Elasticsearch 8.0. By default, when you configure a synchronization or migration task, DTS maps a table from a relational database to an index in Elasticsearch. You can change this mapping when you configure the objects to synchronize or migrate.

    Elasticsearch 7.0 and later

    Elasticsearch

    Relational database

    index

    table

    document

    row

    field

    column

    mapping

    schema

    Versions before Elasticsearch 7.0

    Elasticsearch

    Relational database

    index

    database

    type

    table

    document

    row

    field

    column

    mapping

    schema

Procedure

  1. Navigate to the migration task list page for the destination region using one of the following methods.

    From the DTS console

    1. Log on to the Data Transmission Service (DTS) console.

    2. In the navigation pane on the left, click Data Migration.

    3. In the upper-left corner of the page, select the region where the migration instance is located.

    From the DMS console

    Note

    The actual operations may vary based on the mode and layout of the DMS console. For more information, see Simple mode console and Customize the layout and style of the DMS console.

    1. Log on to the Data Management (DMS) console.

    2. In the top menu bar, choose Data + AI > Data Transmission (DTS) > Data Migration.

    3. To the right of Data Migration Tasks, select the region where the migration instance is located.

  2. Click Create Task to navigate to the task configuration page.

  3. Configure the source and destination databases.

    Warning

    After you select the source and destination instances, we recommend that you carefully read the limits displayed at the top of the page. Otherwise, the task may fail or data inconsistency may occur.

    Section

    Parameter

    Description

    N/A

    Task Name

    DTS automatically generates a task name. We recommend that you specify a descriptive name for easy identification. The name does not need to be unique.

    Source Database

    Select Existing Connection

    • To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.

      Note

      In the DMS console, this parameter is named Select a DMS database instance..

    • If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select PolarDB-X 2.0.

    Access Method

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region where the source PolarDB-X 2.0 instance is located.

    Replicate Data Across Alibaba Cloud Accounts

    This example shows a migration within the same Alibaba Cloud account. Select No.

    Instance ID

    Select the ID of the source PolarDB-X 2.0 instance.

    Database Account

    Enter the database account of the source PolarDB-X 2.0 instance. For information about the required permissions, see Database account permissions.

    Database Password

    Enter the password for the database account.

    Destination Database

    Select Existing Connection

    • To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.

      Note

      In the DMS console, this parameter is named Select a DMS database instance..

    • If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select Elasticsearch.

    Access Method

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region where the destination Elasticsearch instance is located.

    Type

    Select Cluster Edition or Serverless based on your needs.

    Instance ID

    Select the ID of the destination Elasticsearch instance.

    Database Account

    Enter the database account of the destination Elasticsearch instance. The default account is elastic. For information about the required permissions, see Database account permissions.

    Database Password

    Enter the password for the database account.

    Encryption

    Select HTTP or HTTPS as needed.

  4. After you complete the configuration, click Test Connectivity and Proceed at the bottom of the page.

    Note
    • Ensure that the IP address segment of the DTS service is automatically or manually added to the security settings of the source and destination databases to allow access from DTS servers. For more information, see Add DTS server IP addresses to a whitelist.

    • If the source or destination database is a self-managed database (the Access Method is not Alibaba Cloud Instance), you must also click Test Connectivity in the CIDR Blocks of DTS Servers dialog box that appears.

  5. Configure the task objects.

    1. On the Configure Objects page, configure the objects that you want to migrate.

      Parameter

      Description

      Migration Types

      • If you only need to perform a full migration, select both Schema Migration and Full Data Migration.

      • To perform a migration with no downtime, select Schema Migration, Full Data Migration, and Incremental Data Migration.

      Note
      • If you do not select Schema Migration, you must ensure that a database and tables to receive the data exist in the destination database. You can also use the object name mapping feature in the Selected Objects box as needed.

      • If you do not select Incremental Data Migration, do not write new data to the source instance during data migration to ensure data consistency.

      Index Name

      • Table Name

        If you select Table Name, the name of the index created in the destination Elasticsearch instance is the same as the table name. In this example, the index name is order.

      • Database Name_Table Name

        If you select Database Name_Table Name, the name of the index created in the destination Elasticsearch instance is in the Database name_Table name format. In this example, the index name is dtstest_order.

      Processing Mode of Conflicting Tables

      • Precheck and Report Errors: Checks whether tables with the same names exist in the destination database. If no tables with the same names exist, the precheck is passed. If tables with the same names exist, an error is reported during the precheck, and the data migration task does not start.

        Note

        If a table in the destination database has the same name but cannot be easily deleted or renamed, you can change the name of the table in the destination database. For more information, see Object name mapping.

      • Ignore Errors and Proceed: Skips the check for tables with the same names.

        Warning

        Selecting Ignore Errors and Proceed may cause data inconsistency and business risks. For example:

        • If the table schemas are consistent and a record in the destination database has the same primary key value as a record in the source database:

          • During full migration, DTS keeps the record in the destination database. The record from the source database is not migrated.

          • During incremental migration, DTS does not keep the record in the destination database. The record from the source database overwrites the record in the destination database.

        • If the table schemas are inconsistent, only some columns of data may be migrated, or the migration may fail. Proceed with caution.

      Capitalization of Object Names in Destination Instance

      You can configure the case sensitivity policy for the names of databases, tables, and columns in the destination instance. By default, DTS default policy is selected. You can also select a policy that is consistent with the source or destination database. For more information, see Case sensitivity of object names in destination databases.

      Source Objects

      Select one or more objects from the Source Objects section. Click the Rightwards arrow icon and add the objects to the Selected Objects section.

      Note

      The granularity for selecting migration objects is schema, table, and column. If you select only tables or columns as migration objects, other objects such as views, triggers, and stored procedures are not migrated to the destination database.

      Selected Objects

      Note
      • The underscore (_) is the only special character supported in index and type names.

      • If you use the object name mapping feature, the migration of other objects that depend on the renamed object may fail.

      • To set a WHERE condition to filter data, right-click the table that you want to migrate in the Selected Objects section, and set the filter condition in the dialog box that appears. For more information, see Filter task data by using SQL conditions.

      • To select the SQL operations to migrate at the database or table level, right-click an object in the Selected Objects box and select the desired SQL operations in the dialog box that appears.

    2. Click Next: Advanced Settings to configure advanced parameters.

      Parameter

      Description

      Dedicated Cluster for Task Scheduling

      By default, DTS schedules tasks on a shared cluster. You do not need to select one. If you want more stable tasks, you can purchase a dedicated cluster to run DTS migration tasks.

      Retry Time for Failed Connections

      After the migration task starts, if the connection to the source or destination database fails, DTS reports an error and immediately begins to retry the connection. The default retry duration is 720 minutes. You can customize the retry time to a value from 10 to 1440 minutes. We recommend that you set the duration to more than 30 minutes. If DTS reconnects to the source and destination databases within the specified duration, the migration task automatically resumes. Otherwise, the task fails.

      Note
      • For multiple DTS instances that share the same source or destination, the network retry time is determined by the setting of the last created task.

      • Because you are charged for the task during the connection retry period, we recommend that you customize the retry time based on your business needs, or release the DTS instance as soon as possible after the source and destination database instances are released.

      Retry Time for Other Issues

      After the migration task starts, if a non-connectivity issue, such as a DDL or DML execution exception, occurs in the source or destination database, DTS reports an error and immediately begins to retry the operation. The default retry duration is 10 minutes. You can customize the retry time to a value from 1 to 1440 minutes. We recommend that you set the duration to more than 10 minutes. If the related operations succeed within the specified retry duration, the migration task automatically resumes. Otherwise, the task fails.

      Important

      The value of Retry Time for Other Issues must be less than the value of Retry Time for Failed Connections.

      Enable Throttling for Full Data Migration

      During full migration, DTS consumes read and write resources on the source and destination databases, which may increase the database load. If required, you can enable throttling for the full migration task. You can set Queries per second (QPS) to the source database, RPS of Full Data Migration, and Data migration speed for full migration (MB/s) to reduce the load on the destination database.

      Note
      • This configuration item is available only if you select Full Data Migration for Migration Types.

      • You can also adjust the full migration speed after the migration instance is running.

      Enable Throttling for Incremental Data Migration

      If required, you can also choose to set speed limits for the incremental migration task. You can set RPS of Incremental Data Migration and Data migration speed for incremental migration (MB/s) to reduce the load on the destination database.

      Note
      • This configuration item is available only if you select Incremental Data Migration for Migration Types.

      • You can also adjust the incremental migration speed after the migration instance is running.

      Environment Tag

      Optional. Select an environment tag to identify the instance.

      Shard Configuration

      Set the number of primary and replica shards for the index based on the maximum shard configuration of the index in the destination Elasticsearch instance.

      String Index

      The method for indexing strings in the destination Elasticsearch instance.

      • analyzed: Analyzes the strings before indexing them. You must also select a specific analyzer. For more information about analyzer types and their functions, see Analyzers.

      • not analyzed: Indexes the strings with their original values without analysis.

      • no: Does not index the strings.

      Time Zone

      The time zone applied when DTS migrates time-related data types, such as DATETIME and TIMESTAMP, to the destination Elasticsearch instance.

      Note

      If these time-related data types do not require a time zone in the destination instance, you must set the document type (type) for these data types in the destination instance before the migration.

      DOCID

      The DOCID defaults to the primary key of the table. If the table has no primary key, the DOCID is the ID column that Elasticsearch automatically generates.

      Whether to delete SQL operations on heartbeat tables of forward and reverse tasks

      Choose whether DTS writes heartbeat SQL information to the source database while the instance is running.

      • Yes: Does not write heartbeat SQL information to the source database. The DTS instance may display latency.

      • No: Writes heartbeat SQL information to the source database. This may interfere with source database operations like physical backups and cloning.

      Configure ETL

      Choose whether to enable the extract, transform, and load (ETL) feature. For more information, see What is ETL? Valid values:

      Monitoring and Alerting

      Select whether to set alerts and receive alert notifications based on your business needs.

      • No: Does not set an alert.

      • Yes: Configure alerts by setting an alert threshold and an alert contact. If a migration fails or the latency exceeds the threshold, the system sends an alert notification.

    3. After you complete the preceding configurations, click Next: Configure Database and Table Fields at the bottom of the page to set the `_routing` policy and `_id` value for the tables to be migrated in the destination Elasticsearch instance.

      Type

      Description

      Set _routing

      Setting _routing can route and store documents on a specific shard of the destination Elasticsearch instance. For more information, see _routing.

      • Select Yes to use custom columns for routing.

      • Select No to use the _id for routing.

      Note

      If the destination Elasticsearch instance is version 7.x, you must select No.

      Value of _id

      • Primary key column of the table

        A composite primary key is merged into a single column.

      • Business primary key

        If you select Business primary key, you must also set the corresponding Business primary key column.

  6. Save the task and run a precheck.

    • To view the parameters for configuring this instance when you call the API operation, move the pointer over the Next: Save Task Settings and Precheck button and click Preview OpenAPI parameters in the bubble that appears.

    • If you do not need to view or have finished viewing the API parameters, click Next: Save Task Settings and Precheck at the bottom of the page.

    Note
    • Before the migration task starts, DTS performs a precheck. The task starts only after it passes the precheck.

    • If the precheck fails, click View Details next to the failed check item, fix the issue based on the prompt, and then run the precheck again.

    • If a warning is reported during the precheck:

      • For check items that cannot be ignored, click View Details next to the failed item, fix the issue based on the prompt, and then run the precheck again.

      • For check items that can be ignored, you can click Confirm Alert Details, Ignore, OK, and Precheck Again to skip the alert item and run the precheck again. If you choose to ignore a warning, it may cause issues such as data inconsistency and pose risks to your business.

  7. Purchase the instance.

    1. When the Success Rate is 100%, click Next: Purchase Instance.

    2. On the Purchase page, select the link specification for the data migration instance. For more information, see the following table.

      Category

      Parameter

      Description

      New Instance Class

      Resource Group Settings

      Select the resource group to which the instance belongs. The default value is default resource group. For more information, see What is Resource Management?

      Instance Class

      DTS provides migration specifications with different performance levels. The link specification affects the migration speed. You can select a specification based on your business scenario. For more information, see Data migration link specifications.

    3. After the configuration is complete, read and select Data Transmission Service (Pay-as-you-go) Service Terms.

    4. Click Buy and Start. In the OK dialog box that appears, click OK.

      You can view the progress of the migration task on the Data Migration Tasks list page.

      Note
      • If the migration task does not include incremental migration, it stops automatically after the full migration is complete. After the task stops, its Status changes to Completed.

      • If the migration task includes incremental migration, it does not stop automatically. The incremental migration task continues to run. While the incremental migration task is running, the Status of the task is Running.

Check migrated index and data

After the data migration task enters the Running state, connect to the Elasticsearch cluster by using Kibana to verify that the index and data are migrated as expected. For more information, see Log on to the Kibana console.

Note

If the migration results are not as expected, you can delete the index and its data, and reconfigure the data migration task.