Migrate data between buckets in OSS-HDFS

Updated at:

This topic describes how to use Alibaba Cloud Jindo DistCp to migrate data between buckets in an OSS-HDFS service.

Prerequisites

  • You have created an Alibaba Cloud EMR cluster of version EMR-5.6.0 or later, or EMR-3.40.0 or later. For more information, see Create a cluster.

  • Your self-managed cluster built on ECS instances has a Hadoop 2.7+ or 3.x environment and can run MapReduce jobs. Before you migrate data, you must deploy JindoData, which includes JindoSDK and JindoFSx. Download the latest version.

  • You have enabled the OSS-HDFS service and granted the required permissions. For more information, see Enable OSS-HDFS.

Background information

Alibaba Cloud Jindo DistCp is a distributed tool used to copy files within or between large-scale clusters. Jindo DistCp uses MapReduce for file distribution, fault handling, and data recovery. It takes lists of files and folders as input for MapReduce jobs, where each job copies a portion of the files from the source list. Jindo DistCp supports copying data between Hadoop Distributed File System (HDFS) directories, between HDFS and Object Storage Service (OSS), between HDFS and OSS-HDFS, and between OSS-HDFS buckets. It also provides various custom copy parameters and copy policies.

Jindo DistCp provides the following benefits for data migration:

  • High performance: Up to 1.59 times faster than Hadoop DistCp in test scenarios.

  • Rich features: Offers multiple copy methods and scenario-specific optimization strategies.

  • Deep integration with OSS: Supports native OSS features, such as changing a file's storage class to Archive or applying compression.

  • Atomic copy: Guarantees data consistency with a no-rename copy mechanism.

  • Broad compatibility: Serves as a full replacement for Hadoop DistCp and supports Hadoop 2.7+ and Hadoop 3.x.

Step 1: Download the JAR package

To download JindoSDK, visit GitHub.

Step 2: Configure the AccessKey for the OSS-HDFS service

You can configure the AccessKey for the OSS-HDFS service in one of the following ways:

  • Configure the AccessKey in the command

    For example, you can use the --hadoopConf option to configure the AccessKey in the command to migrate data from srcbucket to destbucket in an OSS-HDFS service.

    hadoop jar jindo-distcp-tool-${version}.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=yourkey --hadoopConf fs.oss.accessKeySecret=yoursecret --parallelism 10
  • Configure the AccessKey in a configuration file

    You can configure the `fs.oss.accessKeyId` and `fs.oss.accessKeySecret` parameters for the OSS-HDFS service in the `core-site.xml` file of Hadoop. The following code provides a sample configuration:

    <configuration>
        <property>
            <name>fs.oss.accessKeyId</name>
            <value>LTAI********</value>
        </property>
    
        <property>
            <name>fs.oss.accessKeySecret</name>
            <value>KZo1********</value>
        </property>
    </configuration>

Step 3: Configure the OSS-HDFS service endpoint

You must configure an endpoint to access the OSS-HDFS service. The recommended access path format is oss://<Bucket>.<Endpoint>/<Object>. For example: oss://examplebucket.cn-shanghai.oss-dls.aliyuncs.com/exampleobject.txt. After the configuration is complete, JindoSDK uses the endpoint in the access path to access the corresponding OSS-HDFS service API.

You can also configure the OSS-HDFS service endpoint in other ways. Endpoints that are configured in different ways have different priorities. For more information, see Other ways to configure an endpoint.

Step 4: Perform a full data migration between different buckets in OSS-HDFS

The following example uses Jindo DistCp 4.4.0. You must replace the version number with the one that corresponds to your environment.

Use the same AccessKey to migrate data from Bucket A to Bucket B in the same region

  • Command format

    hadoop jar jindo-distcp-tool-${version}.jar --src oss://bucketname.region.oss-dls.aliyuncs.com/ --dest oss://bucketname.region.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=yourkey --hadoopConf fs.oss.accessKeySecret=yoursecret --parallelism 10

    The following table describes the parameters and options.

    Parameters and options

    Description

    Example

    --src

    The full path of the source bucket in OSS-HDFS from which you want to migrate or copy data.

    oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/

    --dest

    The full path of the destination bucket in OSS-HDFS where you want to store the migrated or copied data.

    oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/

    --hadoopConf

    The AccessKey ID and AccessKey secret used to access the OSS-HDFS service.

    • AccessKey ID

      LTAI********
    • AccessKey secret

      KZo1********

    --parallelism

    The number of concurrent tasks. Adjust this value based on your cluster resources.

    10

  • Example

    You can use the same AccessKey to migrate data from the srcbucket to the destbucket in the China (Hangzhou) region.

    hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --parallelism 10

Migrate data across regions

If the source bucket and the destination bucket reside in different regions, specify the OSS-HDFS endpoint of the corresponding region separately for --src and --dest. The following example migrates data from a bucket in the China (Hangzhou) region to a bucket in the China (Shanghai) region:

hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-shanghai.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --parallelism 10

(Optional) Perform an incremental data migration between different buckets in OSS-HDFS

If you want to copy only the data that was added to the source path after the last full migration, you can perform an incremental migration using the --update option.

For example, you can run the following command to incrementally migrate data from the srcbucket to the destbucket in the China (Hangzhou) region using the same AccessKey.

hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --update --parallelism 10

The --update option also applies to cross-region migration. It copies only the files that changed on the source side since the last migration. This is useful when you need to run multiple migrations or perform incremental backups across regions. The following example incrementally migrates data from a bucket in the China (Hangzhou) region to a bucket in the China (Shanghai) region:

hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-shanghai.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --update --parallelism 10

Cross-region replication (CRR) applicability

The OSS-HDFS feature support list marks cross-region replication (CRR) as not applicable to OSS-HDFS. CRR operates at the object level, while OSS-HDFS metadata is stored under the .dlsdata path. As a result, CRR cannot guarantee consistency at the HDFS metadata layer. To achieve cross-region disaster recovery for OSS-HDFS, use Jindo DistCp to migrate data across regions instead of relying on CRR.

Limits

  • A single Jindo DistCp command supports only one --src and one --dest, and does not support one-to-many synchronization. If you specify multiple --dest values in the same command, the command still completes successfully, but the data is written only to the last specified --dest. The other destinations are not updated.

  • To synchronize the same data to multiple destination buckets, run the DistCp command multiple times and specify a different --dest each time.

  • --disableChecksum: Add this option to the command to disable checksum verification and speed up synchronization.

More information

For information about other Jindo DistCp scenarios, see Jindo DistCp User Guide.