Migrate PolarDB for MySQL to Elasticsearch
Use Data Transmission Service (DTS) to migrate data from PolarDB for MySQL to Elasticsearch.
Prerequisites
-
You have created a destination Elasticsearch cluster. For more information, see Quick start: Create a cluster and retrieve data.
-
The destination Elasticsearch cluster's storage space must exceed that used by the source PolarDB for MySQL instance.
Limitations
|
Type |
Description |
|
Source database limitations |
|
|
Other limitations |
|
|
Other notes |
DTS periodically executes the CREATE DATABASE IF NOT EXISTS |
If a synchronization object of DTS is an index alias in the destination Elasticsearch instance, data inconsistency may occur when the actual index that the alias points to changes. For example, when data is written, the index alias points to Index A. If the alias later points to Index B during a subsequent synchronous delete operation, the data on Index A cannot be deleted. As a result, data deleted from the source can still be queried in the destination.
Billing
|
Migration type |
Instance configuration fee |
Internet traffic fee |
|
Schema migration and full data migration |
Free of charge. |
When the Access Method parameter of the destination database is set to Public IP Address, you are charged for Internet traffic. For more information, see Billing overview. |
|
Incremental data migration |
Charged. For more information, see Billing overview. |
Migration types
-
Schema migration
DTS migrates the schema definitions of the migration objects from the source database to the destination database.
-
Full migration
DTS migrates all historical data of the specified migration objects from the source database to the destination database.
-
Incremental migration
After a full migration is complete, DTS migrates incremental data updates from the source database to the destination database. Incremental migration lets you smoothly migrate data without interrupting your self-managed applications.
SQL operations for incremental migration
Type | SQL statement |
DML | INSERT, UPDATE, and DELETE Note Using an UPDATE statement to remove a column is not a supported operation. |
Permission requirements for database accounts
|
Database |
Permission requirements |
Creation and authorization |
|
PolarDB for MySQL cluster |
Read permission on the objects to be migrated |
Data type mapping
-
A source database and an Elasticsearch instance support different data types that cannot always be mapped directly. During structure initialization, DTS maps data types based on those supported by the target Elasticsearch instance. For more information, see Data type mapping for structure initialization.
NoteDuring the DTS schema migration process, DTS does not set the
dynamicparameter inmapping. The behavior of this parameter depends on the settings of your Elasticsearch instance. If your source data is of the JSON type, you must ensure that for a specific key, its corresponding values have the same data type across all rows in a table. Otherwise, DTS may encounter synchronization issues. For more information, see dynamic. -
The mapping between Elasticsearch and a relational database varies by Elasticsearch version.
ImportantStarting with Elasticsearch 7.0, an index no longer supports multiple types, and types were completely removed in Elasticsearch 8.0. By default, when you configure a synchronization or migration task, DTS maps a table from a relational database to an index in Elasticsearch. You can change this mapping when you configure the objects to synchronize or migrate.
Elasticsearch 7.0 and later
Elasticsearch
Relational database
index
table
document
row
field
column
mapping
schema
Versions before Elasticsearch 7.0
Elasticsearch
Relational database
index
database
type
table
document
row
field
column
mapping
schema
Procedure
-
Navigate to the migration task list page for the destination region using one of the following methods.
From the DTS console
-
Log on to the Data Transmission Service (DTS) console.
-
In the navigation pane on the left, click Data Migration.
-
In the upper-left corner of the page, select the region where the migration instance is located.
From the DMS console
NoteThe actual operations may vary based on the mode and layout of the DMS console. For more information, see Simple mode console and Customize the layout and style of the DMS console.
-
Log on to the Data Management (DMS) console.
-
In the top menu bar, choose .
-
To the right of Data Migration Tasks, select the region where the migration instance is located.
-
-
Click Create Task to navigate to the task configuration page.
-
Configure the source and destination databases.
Section
Parameter
Description
N/A
Task Name
DTS automatically generates a task name. We recommend that you specify a descriptive name for easy identification. The name does not need to be unique.
Source Database
Select Existing Connection
-
To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.
NoteIn the DMS console, this parameter is named Select a DMS database instance..
-
If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.
Database Type
Select PolarDB for MySQL.
Access Method
Select Alibaba Cloud Instance.
Instance Region
Select the region of the source PolarDB for MySQL cluster.
Cross-account
This example migrates data within the same Alibaba Cloud account. Select No.
PolarDB Cluster ID
Select the cluster ID of the source PolarDB for MySQL cluster.
Database Account
Enter the database account of the source PolarDB for MySQL cluster. For information about the required permissions, see Permissions required for the database account.
Database Password
Enter the password for the database account.
Encryption
Select a connection type as needed. For more information about the SSL encryption feature, see Enable SSL encryption.
Destination Database
Select Existing Connection
-
To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.
NoteIn the DMS console, this parameter is named Select a DMS database instance..
-
If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.
Database Type
Select Elasticsearch.
Access Method
Select Alibaba Cloud Instance.
Instance Region
Select the region of the destination Elasticsearch cluster.
Type
Select Cluster or Serverless based on your requirements.
Instance ID
Select the ID of the destination Elasticsearch cluster.
Database Account
Enter the database account for the Elasticsearch cluster. The default account is elastic.
Database Password
Enter the logon password that you set when you created the Elasticsearch cluster.
Encryption
Select HTTP or HTTPS as needed.
-
-
After you complete the configuration, click Test Connectivity and Proceed at the bottom of the page.
Note-
Ensure that the IP address segment of the DTS service is automatically or manually added to the security settings of the source and destination databases to allow access from DTS servers. For more information, see Add DTS server IP addresses to a whitelist.
-
If the source or destination database is a self-managed database (the Access Method is not Alibaba Cloud Instance), you must also click Test Connectivity in the CIDR Blocks of DTS Servers dialog box that appears.
-
-
Configure the task objects.
-
On the Configure Objects page, configure the objects that you want to migrate.
Parameter
Description
Migration Types
-
If you only need to perform a full migration, select both Schema Migration and Full Data Migration.
-
To perform a migration with no downtime, select Schema Migration, Full Data Migration, and Incremental Data Migration.
Note-
If you do not select Schema Migration, you must ensure that a database and tables to receive the data exist in the destination database. You can also use the object name mapping feature in the Selected Objects box as needed.
-
If you do not select Incremental Data Migration, do not write new data to the source instance during data migration to ensure data consistency.
Processing Mode for Existing Destination Tables
-
Precheck and Report Errors: DTS checks for tables in the destination database with the same names as those in the source. The precheck passes if no such tables exist. Otherwise, the precheck fails, and the data migration task does not start.
NoteIf you cannot delete or rename the tables with the same names in the destination database, you can change their names in the destination database. For more information, see object name mapping.
-
Ignore Errors and Proceed: Skips the precheck for tables with the same names.
WarningSelecting Ignore Errors and Proceed may cause data inconsistency and risk service disruption. For example:
-
If the table schemas are consistent and a record in the destination database has the same primary key value as a record in the source database:
-
During full data migration, DTS retains the record in the destination cluster. The record from the source database is not migrated.
-
During incremental data migration, DTS does not retain the record in the destination cluster. The record from the source database overwrites the one in the destination database.
-
-
If the table schemas are inconsistent, data initialization may fail, only partial data may be migrated, or the migration may fail.
-
Index Name
-
Table Name
If you select Table Name, the index created in the destination Elasticsearch instance has the same name as the source table.
-
Database Name_Table Name
If you select Database Name_Table Name, the index created in the destination Elasticsearch instance is named in the DatabaseName_TableName format.
Case Policy for Destination Object Names
Configure case sensitivity for database, table, and column names for the migrated objects in the destination instance. By default, DTS default policy is selected. You can also choose to maintain the same case as the source or destination database. For more information, see Case sensitivity of object names in the destination database.
Source Objects
Select one or more objects from the Source Objects section. Click the
icon and add the objects to the Selected Objects section. NoteYou can select databases or tables as migration objects. If you select tables, other objects such as views, triggers, and stored procedures are not migrated to the destination database.
Selected Objects
-
To change the name of a single migration object in the target instance, right-click the object in the Selected Objects box. For more information, see Map individual schema, table, and column names.
-
To change the names of multiple migration objects in the target instance, click Selected Objects in the upper-right corner of the Batch Edit box. For more information, see Map multiple schema, table, and column names.
Note-
Only underscores (_) are supported as special characters in index names and type names.
-
To filter data with WHERE conditions or specify post-migration details like index, type, and column names, right-click a table in the Selected Objects box. Configure the settings in the dialog box that appears. For more information, see Configure filter conditions.
-
To select SQL operations to migrate at the database or table level, right-click a migration object in the Selected Objects box and select the desired SQL operations in the dialog box that appears. For information about supported operations, see SQL operations supported for incremental data migration.
-
-
Click Next: Advanced Settings to configure advanced parameters.
Parameter
Description
Dedicated Cluster for Task Scheduling
By default, DTS schedules tasks on a shared cluster. You do not need to select one. If you want more stable tasks, you can purchase a dedicated cluster to run DTS migration tasks.
Retry Time for Failed Connections
After the migration task starts, if the connection to the source or destination database fails, DTS reports an error and immediately begins to retry the connection. The default retry duration is 720 minutes. You can customize the retry time to a value from 10 to 1440 minutes. We recommend that you set the duration to more than 30 minutes. If DTS reconnects to the source and destination databases within the specified duration, the migration task automatically resumes. Otherwise, the task fails.
Note-
For multiple DTS instances that share the same source or destination, the network retry time is determined by the setting of the last created task.
-
Because you are charged for the task during the connection retry period, we recommend that you customize the retry time based on your business needs, or release the DTS instance as soon as possible after the source and destination database instances are released.
Retry Time for Other Issues
After the migration task starts, if a non-connectivity issue, such as a DDL or DML execution exception, occurs in the source or destination database, DTS reports an error and immediately begins to retry the operation. The default retry duration is 10 minutes. You can customize the retry time to a value from 1 to 1440 minutes. We recommend that you set the duration to more than 10 minutes. If the related operations succeed within the specified retry duration, the migration task automatically resumes. Otherwise, the task fails.
ImportantThe value of Retry Time for Other Issues must be less than the value of Retry Time for Failed Connections.
Enable Throttling for Full Data Migration
During full migration, DTS consumes read and write resources on the source and destination databases, which may increase the database load. If required, you can enable throttling for the full migration task. You can set Queries per second (QPS) to the source database, RPS of Full Data Migration, and Data migration speed for full migration (MB/s) to reduce the load on the destination database.
Note-
This configuration item is available only if you select Full Data Migration for Migration Types.
-
You can also adjust the full migration speed after the migration instance is running.
Enable Throttling for Incremental Data Migration
If required, you can also choose to set speed limits for the incremental migration task. You can set RPS of Incremental Data Migration and Data migration speed for incremental migration (MB/s) to reduce the load on the destination database.
Note-
This configuration item is available only if you select Incremental Data Migration for Migration Types.
-
You can also adjust the incremental migration speed after the migration instance is running.
Environment Tag
You can select an environment tag to identify the instance based on your requirements. This parameter is not required for this example.
Shard Configuration
Set the number of primary and replica shards for the index based on the maximum shard configuration for an index in the destination Elasticsearch cluster.
String Index
Specifies how to index strings migrated to the destination Elasticsearch cluster.
-
analyzed: Analyzes the strings before indexing. You must also select a specific analyzer. For more information about analyzer types and their functions, see Analyzers.
-
not analyzed: Indexes the strings directly by using their original values without analysis.
-
no: Does not index the strings.
Time Zone
Select the time zone for time-related data types (such as DATETIME and TIMESTAMP) migrated to the destination Elasticsearch instance.
NoteIf the time-related data types in the destination instance do not require a time zone, you must set the document type (type) for these data types in the destination instance before synchronization.
DOCID
By default, the DOCID is the primary key of the table. If the table has no primary key, the DOCID is an ID column automatically generated by Elasticsearch.
Whether to delete SQL operations on heartbeat tables of forward and reverse tasks
Choose whether DTS writes heartbeat SQL information to the source database while the instance is running.
Yes: Does not write heartbeat SQL information to the source database. The DTS instance may display latency.
No: Writes heartbeat SQL information to the source database. This may interfere with source database operations like physical backups and cloning.
Configure ETL
Based on your business needs, select whether to configure the ETL feature to process data.
-
Yes: Configures the ETL feature. You must also enter data processing statements in the text box.
-
No: Does not configure the ETL feature.
Monitoring and Alerting
Select whether to set alerts and receive alert notifications based on your business needs.
-
No: Does not set an alert.
-
Yes: Configure alerts by setting an alert threshold and an alert contact. If a migration fails or the latency exceeds the threshold, the system sends an alert notification.
-
-
After completing the preceding settings, click Next: Configure Database and Table Fields to set the _routing policy and _id value for the tables to be migrated to the destination Elasticsearch instance.
Parameter
Description
Set _routing
The _routing parameter routes a document to a specific shard in the destination Elasticsearch instance. For more information, see _routing.
-
Select Yes to specify custom columns for routing.
-
Select No to use the _id value for routing.
NoteIf the destination Elasticsearch instance is version 7.x, select No.
Value of _id
-
Primary key column
A composite primary key is merged into a single column.
-
Business key
If you select Business key, you must also specify the business key column.
-
-
-
Save the task and run a precheck.
-
To view the parameters for configuring this instance when you call the API operation, move the pointer over the Next: Save Task Settings and Precheck button and click Preview OpenAPI parameters in the bubble that appears.
-
If you do not need to view or have finished viewing the API parameters, click Next: Save Task Settings and Precheck at the bottom of the page.
Note-
Before the migration task starts, DTS performs a precheck. The task starts only after it passes the precheck.
-
If the precheck fails, click View Details next to the failed check item, fix the issue based on the prompt, and then run the precheck again.
-
If a warning is reported during the precheck:
-
For check items that cannot be ignored, click View Details next to the failed item, fix the issue based on the prompt, and then run the precheck again.
-
For check items that can be ignored, you can click Confirm Alert Details, Ignore, OK, and Precheck Again to skip the alert item and run the precheck again. If you choose to ignore a warning, it may cause issues such as data inconsistency and pose risks to your business.
-
-
-
Purchase the instance.
-
When the Success Rate is 100%, click Next: Purchase Instance.
-
On the Purchase page, select the link specification for the data migration instance. For more information, see the following table.
Category
Parameter
Description
New Instance Class
Resource Group Settings
Select the resource group to which the instance belongs. The default value is default resource group. For more information, see What is Resource Management?
Instance Class
DTS provides migration specifications with different performance levels. The link specification affects the migration speed. You can select a specification based on your business scenario. For more information, see Data migration link specifications.
-
After the configuration is complete, read and select Data Transmission Service (Pay-as-you-go) Service Terms.
-
Click Buy and Start. In the OK dialog box that appears, click OK.
You can view the progress of the migration task on the Data Migration Tasks list page.
Note-
If the migration task does not include incremental migration, it stops automatically after the full migration is complete. After the task stops, its Status changes to Completed.
-
If the migration task includes incremental migration, it does not stop automatically. The incremental migration task continues to run. While the incremental migration task is running, the Status of the task is Running.
-
-