Migrate data between buckets in OSS-HDFS
This topic describes how to use Alibaba Cloud Jindo DistCp to migrate data between buckets in an OSS-HDFS service.
Prerequisites
-
You have created an Alibaba Cloud EMR cluster of version EMR-5.6.0 or later, or EMR-3.40.0 or later. For more information, see Create a cluster.
-
Your self-managed cluster built on ECS instances has a Hadoop 2.7+ or 3.x environment and can run MapReduce jobs. Before you migrate data, you must deploy JindoData, which includes JindoSDK and JindoFSx. Download the latest version.
-
You have enabled the OSS-HDFS service and granted the required permissions. For more information, see Enable OSS-HDFS.
Background information
Alibaba Cloud Jindo DistCp is a distributed tool used to copy files within or between large-scale clusters. Jindo DistCp uses MapReduce for file distribution, fault handling, and data recovery. It takes lists of files and folders as input for MapReduce jobs, where each job copies a portion of the files from the source list. Jindo DistCp supports copying data between Hadoop Distributed File System (HDFS) directories, between HDFS and Object Storage Service (OSS), between HDFS and OSS-HDFS, and between OSS-HDFS buckets. It also provides various custom copy parameters and copy policies.
Jindo DistCp provides the following benefits for data migration:
-
High performance: Up to 1.59 times faster than Hadoop DistCp in test scenarios.
-
Rich features: Offers multiple copy methods and scenario-specific optimization strategies.
-
Deep integration with OSS: Supports native OSS features, such as changing a file's storage class to Archive or applying compression.
-
Atomic copy: Guarantees data consistency with a no-rename copy mechanism.
-
Broad compatibility: Serves as a full replacement for Hadoop DistCp and supports Hadoop 2.7+ and Hadoop 3.x.
Step 1: Download the JAR package
To download JindoSDK, visit GitHub.
Step 2: Configure the AccessKey for the OSS-HDFS service
You can configure the AccessKey for the OSS-HDFS service in one of the following ways:
-
Configure the AccessKey in the command
hadoop jar jindo-distcp-tool-${version}.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=yourkey --hadoopConf fs.oss.accessKeySecret=yoursecret --parallelism 10 -
Configure the AccessKey in a configuration file
You can configure the `fs.oss.accessKeyId` and `fs.oss.accessKeySecret` parameters for the OSS-HDFS service in the `core-site.xml` file of Hadoop. The following code provides a sample configuration:
<configuration> <property> <name>fs.oss.accessKeyId</name> <value>LTAI********</value> </property> <property> <name>fs.oss.accessKeySecret</name> <value>KZo1********</value> </property> </configuration>
Step 3: Configure the OSS-HDFS service endpoint
You must configure an endpoint to access the OSS-HDFS service. The recommended access path format is oss://<Bucket>.<Endpoint>/<Object>. For example: oss://examplebucket.cn-shanghai.oss-dls.aliyuncs.com/exampleobject.txt. After the configuration is complete, JindoSDK uses the endpoint in the access path to access the corresponding OSS-HDFS service API.
You can also configure the OSS-HDFS service endpoint in other ways. Endpoints that are configured in different ways have different priorities. For more information, see Other ways to configure an endpoint.
Step 4: Perform a full data migration between different buckets in OSS-HDFS
The following example uses Jindo DistCp 4.4.0. You must replace the version number with the one that corresponds to your environment.
Use the same AccessKey to migrate data from Bucket A to Bucket B in the same region
-
Command format
hadoop jar jindo-distcp-tool-${version}.jar --src oss://bucketname.region.oss-dls.aliyuncs.com/ --dest oss://bucketname.region.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=yourkey --hadoopConf fs.oss.accessKeySecret=yoursecret --parallelism 10The following table describes the parameters and options.
Parameters and options
Description
Example
--src
The full path of the source bucket in OSS-HDFS from which you want to migrate or copy data.
oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/
--dest
The full path of the destination bucket in OSS-HDFS where you want to store the migrated or copied data.
oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/
--hadoopConf
The AccessKey ID and AccessKey secret used to access the OSS-HDFS service.
-
AccessKey ID
LTAI******** -
AccessKey secret
KZo1********
--parallelism
The number of concurrent tasks. Adjust this value based on your cluster resources.
10
-
-
Example
You can use the same AccessKey to migrate data from the srcbucket to the destbucket in the China (Hangzhou) region.
hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --parallelism 10
Migrate data across regions
If the source bucket and the destination bucket reside in different regions, specify the OSS-HDFS endpoint of the corresponding region separately for --src and --dest. The following example migrates data from a bucket in the China (Hangzhou) region to a bucket in the China (Shanghai) region:
hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-shanghai.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --parallelism 10
(Optional) Perform an incremental data migration between different buckets in OSS-HDFS
If you want to copy only the data that was added to the source path after the last full migration, you can perform an incremental migration using the --update option.
For example, you can run the following command to incrementally migrate data from the srcbucket to the destbucket in the China (Hangzhou) region using the same AccessKey.
hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --update --parallelism 10
The --update option also applies to cross-region migration. It copies only the files that changed on the source side since the last migration. This is useful when you need to run multiple migrations or perform incremental backups across regions. The following example incrementally migrates data from a bucket in the China (Hangzhou) region to a bucket in the China (Shanghai) region:
hadoop jar jindo-distcp-tool-4.4.0.jar --src oss://srcbucket.cn-hangzhou.oss-dls.aliyuncs.com/ --dest oss://destbucket.cn-shanghai.oss-dls.aliyuncs.com/ --hadoopConf fs.oss.accessKeyId=LTAI******** --hadoopConf fs.oss.accessKeySecret=KZo1******** --update --parallelism 10
Cross-region replication (CRR) applicability
The OSS-HDFS feature support list marks cross-region replication (CRR) as not applicable to OSS-HDFS. CRR operates at the object level, while OSS-HDFS metadata is stored under the .dlsdata path. As a result, CRR cannot guarantee consistency at the HDFS metadata layer. To achieve cross-region disaster recovery for OSS-HDFS, use Jindo DistCp to migrate data across regions instead of relying on CRR.
Limits
-
A single Jindo DistCp command supports only one
--srcand one--dest, and does not support one-to-many synchronization. If you specify multiple--destvalues in the same command, the command still completes successfully, but the data is written only to the last specified--dest. The other destinations are not updated. -
To synchronize the same data to multiple destination buckets, run the DistCp command multiple times and specify a different
--desteach time. -
--disableChecksum: Add this option to the command to disable checksum verification and speed up synchronization.
More information
For information about other Jindo DistCp scenarios, see Jindo DistCp User Guide.