Migrate ApsaraDB for MongoDB to Lindorm

更新时间:
复制 MD 格式

This topic describes how to use Data Transmission Service (DTS) to migrate data from an ApsaraDB for MongoDB instance (replica set or sharded cluster) to a Lindorm instance.

Prerequisites

Limitations

Type

Description

Limits on the source database

  • The source database server must have sufficient egress bandwidth. Insufficient bandwidth affects data migration speed.

  • The collections that you want to migrate must have a primary key or a unique constraint, and the fields must have unique values. Otherwise, duplicate data may occur in the destination database.

  • If you migrate data at the collection level and need to edit the collections, such as by mapping collection names, a single migration task can migrate a maximum of 1,000 collections. If you exceed this limit, the task submission fails. In this case, split the collections into multiple batches and configure a separate task for each batch.

  • If the source database is an Azure Cosmos DB for MongoDB or an Amazon DocumentDB elastic cluster, only full data migration is supported.

  • To perform incremental data migration:

    The source database must have the oplog enabled with at least seven days of retention. Alternatively, change streams must be enabled, and DTS must be able to subscribe to data changes from the source database within the last seven days by using change streams. If these requirements are not met, the migration task may fail because it cannot obtain data changes from the source. In extreme cases, this can lead to data inconsistency or loss. Issues arising from this are not covered by the DTS Service Level Agreement (SLA).

    Important
    • We recommend using the oplog to obtain data changes from the source database.

    • Only MongoDB 4.0 and later versions support obtaining data changes through change streams.

    • If the source database is an Amazon DocumentDB (non-elastic) cluster, you must manually enable change streams. When you configure the task, set the Migration Method to ChangeStream and the Architecture to Sharded Cluster.

  • If the source ApsaraDB for MongoDB instance has a sharded cluster architecture, the _id field in the collections to be migrated must be unique. Otherwise, data inconsistency may occur.

  • If the source ApsaraDB for MongoDB instance has a sharded cluster architecture, the number of source mongos nodes cannot exceed 10.

  • If the source instance is a self-managed MongoDB with a sharded cluster architecture:

    • The Access Method supports only Public IP Address, Express Connect, VPN Gateway, or Smart Access Gateway, and Cloud Enterprise Network (CEN).

    • If the MongoDB version is 8.0 or later and the Migration Method is Oplog, you must ensure that the shard account used by the migration task has the directShardOperations permission. You can grant this permission by running the following command: db.adminCommand({ grantRolesToUser: "username", roles: [{ role: "directShardOperations", db: "admin"}]})

      Note

      Replace username in the command with the shard account used by the migration task.

    • If the Migration Method is Oplog and the task includes full migration, you must ensure that the mongos account of the source MongoDB sharded cluster has the permission to run the db.runCommand({"balancerStatus":1}) command. DTS uses this command during the precheck phase to verify that the source Balancer is disabled.

  • DTS does not support connecting to a MongoDB database by using an SRV record.

  • Operational limits on the source database:

    • During the full migration phase, do not perform schema changes on databases or collections, including updating data in arrays. Such changes can cause the migration task to fail or lead to data inconsistency between the source and destination databases.

    • If you perform only a full data migration, do not write new data to the source instance. Otherwise, data will be inconsistent between the source and destination databases.

    • If the source MongoDB instance is a sharded cluster, do not run commands that change data distribution for the objects to be migrated on the source database while the migration task is running, such as shardCollection, reshardCollection, unshardCollection, moveCollection, or movePrimary. Otherwise, data inconsistency may occur.

  • If a collection to be migrated contains a Time-to-Live (TTL) index, data inconsistency may occur or instance latency may increase.

  • If the source database is a MongoDB that uses the sharded cluster architecture and the source Balancer is rebalancing data, the source instance may experience latency.

Other limits

  • DTS does not support migrating data from the admin, config, or local databases.

  • The collections in the destination Lindorm instance cannot have fields named _id and _value. Otherwise, the migration will fail.

  • If your incremental migration includes UPDATE or DELETE operations, the following limits apply:

    • If the wide table is created by using Lindorm SQL, you must add a non-primary key column named _mongo_id_ when you create the table. The data type of this column depends on the data type of the _id field in MongoDB. You must also create a secondary index for this column.

    • If the wide table is created by using the HBase API, you must add a non-primary key column named _mongo_id_ to the f column family when you create the table. The data type of this column depends on the data type of the _id field in MongoDB. You must also create a secondary index for this column. If you use the new column with the ETL feature, ensure that no duplicate data exists in Lindorm.

  • Transaction information is not preserved. Transactions in the source database become single records in the destination database.

  • The data to be synchronized in the Lindorm instance must meet the requirements specified in Request limits. Otherwise, the task fails.

  • Before migrating data, evaluate the performance of both the source and destination databases. Perform the data migration during off-peak hours. A full data migration consumes read and write resources on both databases, which increases the load on them.

  • During full data migration, DTS performs concurrent INSERT operations. This can cause fragmentation in the destination collections. As a result, the destination collections will use more storage space than the source collections.

  • Verify that the migration precision that DTS uses for FLOAT or DOUBLE values meets your business requirements. DTS reads values of these types by using ROUND(COLUMN,PRECISION). If you do not specify the precision, DTS uses a default precision of 38 digits for FLOAT and 308 digits for DOUBLE.

  • DTS attempts to resume failed migration tasks within seven days. Therefore, before you switch workloads to the destination instance, you must stop or release the task, or use the revoke command to revoke write permissions from the account that DTS uses to access the destination instance. This prevents the task from automatically resuming and overwriting data in the destination instance.

  • DTS calculates the latency for an incremental migration by comparing the timestamp of the latest record migrated to the destination database with the current timestamp. If the source database has no recent updates, the reported latency may be inaccurate. If the task displays excessive latency, you can perform a write operation on the source database to get an accurate latency reading.

  • Time series collections, introduced in MongoDB 5.0, are not supported for migration.

  • If a task fails, DTS support staff will attempt to restore it within eight hours. During restoration, they may restart the task or adjust its parameters.

    Note

    Only DTS task parameters are modified—not database parameters. Parameters that may be adjusted include those listed in Modify instance parameters.

Billing

Migration type

Task configuration fee

Data transfer fee

full data migration

Free.

This tutorial is free. However, a data transfer fee applies if the Access Method for the destination database is Public IP Address.

incremental data migration

Fees apply. For details, see billing overview.

Migration types

Migration type

Description

Full data migration

Migrates all existing data from the source ApsaraDB for MongoDB instance to the destination Lindorm instance.

Note

Supports full data migration for databases and collections.

Incremental data migration

After a full data migration, migrates incremental updates from the source ApsaraDB for MongoDB instance to the destination Lindorm instance.

Note
  • Supports only insert, update, and delete operations on documents in a collection.

  • For incremental document updates, only changes made by using the $set command are replicated.

Database account permissions

Database

Full data migration

Incremental data migration

Actions

Source ApsaraDB for MongoDB

The read permission on the databases to be migrated.

The read permission on the databases to be migrated, the admin database, and the local database.

Manage MongoDB database account permissions

Target Lindorm

The read and write permissions on the Lindorm instance.

User management

Note

If you use ChangeStream as the incremental migration method, the source database account requires instance-wide Change Streams read permissions (such as readAnyDatabase). If the source is an ApsaraDB for MongoDB instance with a custom account, you must also grant the account read permission on the admin database. For details, see Permissions of the root account specified during instance creation.

Procedure

This procedure uses a wide table, created in Lindorm with Lindorm SQL, as an example destination database.

  1. Navigate to the migration task list page for the destination region using one of the following methods.

    From the DTS console

    1. Log on to the Data Transmission Service (DTS) console.

    2. In the navigation pane on the left, click Data Migration.

    3. In the upper-left corner of the page, select the region where the migration instance is located.

    From the DMS console

    Note

    The actual operations may vary based on the mode and layout of the DMS console. For more information, see Simple mode console and Customize the layout and style of the DMS console.

    1. Log on to the Data Management (DMS) console.

    2. In the top menu bar, choose Data + AI > Data Transmission (DTS) > Data Migration.

    3. To the right of Data Migration Tasks, select the region where the migration instance is located.

  2. Click Create Task to navigate to the task configuration page.

  3. Configure the source and destination databases.

    Section

    Parameter

    Description

    N/A

    Task Name

    DTS automatically generates a task name. We recommend that you specify a descriptive name for easy identification. The name does not need to be unique.

    Source Database

    Select Existing Connection

    • To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.

      Note

      In the DMS console, this parameter is named Select a DMS database instance..

    • If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select MongoDB.

    Connection Type

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region where the source ApsaraDB for MongoDB instance resides.

    Replicate Data Across Alibaba Cloud Accounts

    In this example, data is migrated within the same Alibaba Cloud account. Select No.

    Architecture

    In this example, Replica Set is selected.

    Note

    If you set Architecture to Sharded Cluster, you must also specify Shard account and Shard password.

    Migration Method

    Select a method for incremental data migration based on your requirements.

    • Oplog (Recommended):

      This option is available if an oplog is enabled for the source database.

      Note

      An oplog is enabled by default for both self-managed MongoDB databases and ApsaraDB for MongoDB instances. This method offers lower latency for incremental data migration due to faster log pulling. Therefore, we recommend selecting Oplog.

    • ChangeStream: This option is available if Change Streams are enabled for the source database.

      Note
      • If the source database is an Amazon DocumentDB instance (non-elastic cluster), you can only select ChangeStream.

      • If you set Architecture to Sharded Cluster for the source database, you do not need to enter a Shard account or Shard password.

    Instance ID

    Select the ID of the source ApsaraDB for MongoDB instance.

    Authentication Database

    Enter the name of the database to which the database account of the source ApsaraDB for MongoDB instance belongs. The default value is admin.

    Database Account

    Enter the database account of the source ApsaraDB for MongoDB instance.

    Database Password

    Enter the password for the specified database account.

    Encryption

    DTS supports three connection methods: Non-encrypted, SSL-encrypted, and Mongo Atlas SSL. The options for Encryption vary based on the selected Access Method and Architecture. The options displayed in the console prevail.

    Note
    • A MongoDB database where the Architecture is Sharded Cluster and the Migration Method is Oplog does not support SSL-encrypted.

    • If the source is a self-managed MongoDB database (Access Method is not Alibaba Cloud Instance) with a Replica Set architecture, and you select SSL-encrypted, DTS also allows you to upload a CA certificate to verify the connection.

    Destination Database

    Select Existing Connection

    • To use a database instance that has been added to the system (created or saved), select the desired database instance from the drop-down list. The database information below will be automatically configured.

      Note

      In the DMS console, this parameter is named Select a DMS database instance..

    • If you have not registered the database instance with the system, or do not need to use a registered instance, manually configure the database information below.

    Database Type

    Select Lindorm.

    Connection Type

    Select Alibaba Cloud Instance.

    Instance Region

    Select the region where the destination Lindorm instance resides.

    Instance ID

    Select the ID of the destination Lindorm instance.

    Database Account

    Enter the database account of the destination Lindorm instance.

    Database Password

    Enter the password for the specified database account.

  4. After you complete the configuration, click Test Connectivity and Proceed at the bottom of the page.

    Note
    • Ensure that the IP address segment of the DTS service is automatically or manually added to the security settings of the source and destination databases to allow access from DTS servers. For more information, see Add DTS server IP addresses to a whitelist.

    • If the source or destination database is a self-managed database (the Access Method is not Alibaba Cloud Instance), you must also click Test Connectivity in the CIDR Blocks of DTS Servers dialog box that appears.

  5. Configure the task objects.

    1. On the Configure Objects page, configure the objects that you want to migrate.

      Parameter

      Description

      Migration Types

      • If you only need to perform a full migration, select Full Data Migration.

      • To perform a migration with no downtime, select both Full Data Migration and Incremental Data Migration.

      Note

      If you do not select Incremental Data Migration, do not write new data to the source instance during data migration to ensure data consistency.

      Processing Mode of Conflicting Tables

      You can retain the default settings.

      Capitalization of Object Names in Destination Instance

      You can configure the case sensitivity policy for the names of migrated objects, such as databases and collections, in the destination instance. By default, DTS default policy is selected. You can also choose to keep the case sensitivity consistent with the default policy of the source or destination database. For more information, see Case sensitivity of object names in the destination database.

      Source Objects

      In the Source Objects pane, click the collections that you want to migrate and click the 向右小箭头 icon to move them to the Selected Objects pane.

      Selected Objects

      If the destination wide table was created using Lindorm SQL, you must add new columns for data migration. DTS will not migrate unconfigured columns.

      1. Edit the database name mapping.

        1. In the Selected Objects pane, right-click the database that contains the collection to migrate.

        2. Change the Schema Name to the name of the destination database in Lindorm.

          image.png

        3. Optional: In the Select DDL and DML Operations to Be Synchronized section, select the operations for incremental migration.

        4. Click OK.

      2. Edit the table name mapping.

        1. In the Selected Objects pane, right-click the collection to migrate.

        2. Change the Table Name to the name of the destination table in Lindorm.

          image.png

        3. Optional: Specify filter conditions. For more information, see Set filter conditions.

        4. Optional: In the Select DDL and DML Operations to Be Synchronized section, select the operations for incremental migration.

      3. Configure the fields to migrate from MongoDB.

        By default, DTS maps the data of the collection to migrate and configures an expression in the Parameter Value column. You need to check whether the expression meets your requirements and configure parameters such as Column Name, Type, Length, and Precision.

        1. In the Parameter Value column, view the field name for the row in MongoDB in the bson_value() expression.

          The string in "" specifies the field name in MongoDB. For example, if the expression is bson_value("age"), the data in this row corresponds to the age field in MongoDB.

        2. Optional: Delete the fields that you do not want to migrate.

          Note

          To delete a field that you do not want to migrate, click the image icon in the field's row.

        3. Configure the fields to migrate.

          Take further action based on whether the bson_value() expression meets your requirements.

          Matching expressions

          1. Enter a Column Name.

            Note

            Enter the name of the destination column in the Lindorm table.

            • If the destination table is created by using SQL, set Column Name to the name of the destination column in the Lindorm table.

            • If the destination table is created by using the HBase API and you need to add new columns, you must add column mappings before you modify column names. For more information, see Example of adding column mappings for a table created by calling the Apache HBase API. Set Column Name based on the following rules:

              • If the column is a primary key, set the name to ROW.

              • If the column is not a primary key, use the column family:column name format. Example: person:name.

          2. Select a data Type for the column.

            Important

            Make sure that the data type of the destination table is compatible with the data type in the source MongoDB database.

          3. Optional: Configure the Length and Precision of the column data.

          4. Repeat the preceding steps to map each required field.

          Custom expressions

          Note

          For example, fields with hierarchical relationships (parent-child structures).

          1. In the Actions column, click the image icon in the row of the field.

          2. Click + Add Column. image

          3. Configure the Column Name, Type, Length, and Precision.

          4. In the text box under Parameter Value, enter the bson_value() expression. For more information, see Value Configuration Example.

            Important
            • The primary key column of the target table must be assigned the value bson_value("_id").

            • When you configure the bson_value() expression, you must specify the path to the lowest-level subfield. Otherwise, data loss or task failure may occur.

          5. Repeat the preceding steps to map each required field.

      4. Click OK.

    2. Click Next: Advanced Settings to configure advanced parameters.

      Parameter

      Description

      Dedicated Cluster for Task Scheduling

      By default, DTS schedules tasks on a shared cluster. You do not need to select one. If you want more stable tasks, you can purchase a dedicated cluster to run DTS migration tasks.

      Retry Time for Failed Connections

      After the migration task starts, if the connection to the source or destination database fails, DTS reports an error and immediately begins to retry the connection. The default retry duration is 720 minutes. You can customize the retry time to a value from 10 to 1440 minutes. We recommend that you set the duration to more than 30 minutes. If DTS reconnects to the source and destination databases within the specified duration, the migration task automatically resumes. Otherwise, the task fails.

      Note
      • For multiple DTS instances that share the same source or destination, the network retry time is determined by the setting of the last created task.

      • Because you are charged for the task during the connection retry period, we recommend that you customize the retry time based on your business needs, or release the DTS instance as soon as possible after the source and destination database instances are released.

      Retry Time for Other Issues

      After the migration task starts, if a non-connectivity issue, such as a DDL or DML execution exception, occurs in the source or destination database, DTS reports an error and immediately begins to retry the operation. The default retry duration is 10 minutes. You can customize the retry time to a value from 1 to 1440 minutes. We recommend that you set the duration to more than 10 minutes. If the related operations succeed within the specified retry duration, the migration task automatically resumes. Otherwise, the task fails.

      Important

      The value of Retry Time for Other Issues must be less than the value of Retry Time for Failed Connections.

      Enable Throttling for Full Data Migration

      During full migration, DTS consumes read and write resources on the source and destination databases, which may increase the database load. If required, you can enable throttling for the full migration task. You can set Queries per second (QPS) to the source database, RPS of Full Data Migration, and Data migration speed for full migration (MB/s) to reduce the load on the destination database.

      Note
      • This configuration item is available only if you select Full Data Migration for Migration Types.

      • You can also adjust the full migration speed after the migration instance is running.

      Only one data type for primary key _id in a table of the data to be synchronized

      In the data to be migrated, is the data type of the primary key _id uniform within a single collection?

      Important
      • Select an option based on your requirements. Otherwise, data loss may occur.

      • This parameter is available only if you select Full Data Migration for Migration Types.

      • Yes: The data type is unique. During full data migration, DTS does not scan the data types of primary keys in the source data. For a single collection, DTS migrates only the data corresponding to one primary key data type.

      • No: The data type is not unique. During full data migration, DTS scans the data types of primary keys in the source data and migrates all data.

      Enable Throttling for Incremental Data Migration

      If required, you can also choose to set speed limits for the incremental migration task. You can set RPS of Incremental Data Migration and Data migration speed for incremental migration (MB/s) to reduce the load on the destination database.

      Note
      • This configuration item is available only if you select Incremental Data Migration for Migration Types.

      • You can also adjust the incremental migration speed after the migration instance is running.

      Environment Tag

      You can select an environment tag to identify the instance based on your business requirements. In this example, you do not need to select a tag.

      Configure ETL

      Based on your business needs, select whether to configure the ETL feature to process data.

      • Yes: Configures the ETL feature. You must also enter data processing statements in the text box.

      • No: Does not configure the ETL feature.

      Note

      If the destination table is created by using the HBase API, note the following:

      • The ETL syntax includes columns to configure and columns to exclude. During migration, all top-level fields of the MongoDB documents for which ETL is configured are stored in the default column family f of the HBase table. The following example shows how to write all elements except for the top-level elements _id and name as dynamic columns to the destination table. For more information, see Example of configuring an ETL task for a table created by calling the Apache HBase API.

        script:e_expand_bson_value("*", "_id,name")
      • If you need to use both the new column and ETL features, make sure that no duplicate data exists in the Lindorm instance.

      • Columns for which neither the new column nor the ETL feature is configured are not migrated to the destination database.

      Monitoring and Alerting

      Select whether to set alerts and receive alert notifications based on your business needs.

      • No: Does not set an alert.

      • Yes: Configure alerts by setting an alert threshold and an alert contact. If a migration fails or the latency exceeds the threshold, the system sends an alert notification.

  6. Save the task and run a precheck.

    • To view the parameters for configuring this instance when you call the API operation, move the pointer over the Next: Save Task Settings and Precheck button and click Preview OpenAPI parameters in the bubble that appears.

    • If you do not need to view or have finished viewing the API parameters, click Next: Save Task Settings and Precheck at the bottom of the page.

    Note
    • Before the migration task starts, DTS performs a precheck. The task starts only after it passes the precheck.

    • If the precheck fails, click View Details next to the failed check item, fix the issue based on the prompt, and then run the precheck again.

    • If a warning is reported during the precheck:

      • For check items that cannot be ignored, click View Details next to the failed item, fix the issue based on the prompt, and then run the precheck again.

      • For check items that can be ignored, you can click Confirm Alert Details, Ignore, OK, and Precheck Again to skip the alert item and run the precheck again. If you choose to ignore a warning, it may cause issues such as data inconsistency and pose risks to your business.

  7. Purchase the instance.

    1. When the Success Rate reaches 100%, click Next: Purchase Instance.

    2. On the Purchase page, select the link specification for the data migration instance. For more information, see the following table.

      Category

      Parameter

      Description

      New Instance Class

      Resource Group Settings

      Select the resource group to which the instance belongs. The default value is default resource group. For more information, see What is Resource Management?

      Instance Class

      DTS provides migration specifications with different performance levels. The link specification affects the migration speed. You can select a specification based on your business scenario. For more information, see Data migration link specifications.

    3. After the configuration is complete, read and select Data Transmission Service (Pay-as-you-go) Service Terms.

    4. Click Buy and Start. In the OK dialog box that appears, click OK.

      You can view the progress of the migration task on the Data Migration Tasks list page.

      Note
      • If the migration task does not include incremental migration, it stops automatically after the full migration is complete. After the task stops, its Status changes to Completed.

      • If the migration task includes incremental migration, it does not stop automatically. The incremental migration task continues to run. While the incremental migration task is running, the Status of the task is Running.

HBase table column mapping

This example shows the commands to run in the SQL Shell.

Note

This feature requires Lindorm 2.4.0 or later.

  1. Add a column mapping to the HBase table.

    ALTER TABLE test MAP DYNAMIC COLUMN f:_mongo_id_ HSTRING/HINT/..., person:name HSTRING, person:age HINT;
  2. Create a secondary index on the HBase table.

    CREATE INDEX idx ON test(f:_mongo_id_);

HBase table migration (ETL)

MongoDB document

{
  "_id" : 0,
  "person" : {
    "name" : "cindy0",
    "age" : 0,
    "student" : true
  }
}

ETL statement

script:e_expand_bson_value("*", "_id")

Migration result

迁移结果

Assignment configuration example

Source ApsaraDB for MongoDB structure

{
  "_id":"62cd344c85c1ea6a2a9f****",
  "person":{
    "name":"neo",
    "age":"26",
    "sex":"male"
  }
}

Destination Lindorm table schema

Parameter

Type

id

STRING

person_name

STRING

person_age

BIGINT

Additional column configuration

Important

Configure the bson_value() expression according to the data hierarchy to prevent data loss or task failure. For example, if you configure the expression as bson_value("person"), Data Transmission Service (DTS) cannot write incremental changes from the sub-fields of the source person object, such as name, age, and sex, to the destination.

Parameter

Type

Value

id

STRING

bson_value("_id")

person_name

STRING

bson_value("person","name")

person_age

BIGINT

bson_value("person","age")