The multipart upload feature of Object Storage Service (OSS) lets you split a large object into multiple parts. After these parts are uploaded, you can call the CompleteMultipartUpload operation to combine the parts into a complete object.
Notes
The sample code in this topic uses the China (Hangzhou) region (
cn-hangzhou) and its public endpoint by default. If you want to access OSS from other Alibaba Cloud products in the same region, use an internal endpoint. For more information about the regions and endpoints supported by OSS, see Regions and endpoints.To perform a multipart upload, you must have the
oss:PutObjectpermission. For more information, see Grant a RAM user custom permissions.
Multipart upload process
A multipart upload consists of the following three steps:
Initialize a multipart upload event.
Call the Client.InitiateMultipartUpload method to obtain a globally unique upload ID from OSS.
Upload parts.
Call the Client.UploadPart method to upload part data.
NoteFor the same upload ID, the part number identifies the part and its relative position in the complete object. If you upload a new part with an existing part number, the existing part on OSS is overwritten.
OSS includes the MD5 hash of the received part data in the ETag header of the response.
OSS calculates the MD5 hash of the uploaded data and compares it with the MD5 hash calculated by the software development kit (SDK). If the two MD5 hashes are different, the InvalidDigest error code is returned.
Complete the multipart upload.
After all parts are uploaded, you can call the Client.CompleteMultipartUpload method to merge all parts into a complete object.
Sample code
The following sample code shows how to split a large local file into multiple parts, concurrently upload the parts to a bucket, and then merge the parts into a complete object.
import os
import argparse
import alibabacloud_oss_v2 as oss
# Create a command-line argument parser for the multipart upload sample.
parser = argparse.ArgumentParser(description="multipart upload sample")
# Add the required --region command-line argument, which specifies the region where the bucket is located.
parser.add_argument('--region', help='The region in which the bucket is located.', required=True)
# Add the required --bucket command-line argument, which specifies the name of the bucket.
parser.add_argument('--bucket', help='The name of the bucket.', required=True)
# Add the optional --endpoint command-line argument, which specifies the domain name that other services can use to access OSS.
parser.add_argument('--endpoint', help='The domain names that other services can use to access OSS')
# Add the required --key command-line argument, which specifies the name of the object.
parser.add_argument('--key', help='The name of the object.', required=True)
# Add the required --file_path command-line argument, which specifies the path of the file to upload.
parser.add_argument('--file_path', help='The path of Upload file.', required=True)
def main():
# Parse the command-line arguments.
args = parser.parse_args()
# Load credentials from environment variables for identity verification.
credentials_provider = oss.credentials.EnvironmentVariableCredentialsProvider()
# Use the default configurations of the SDK and set the credentials provider.
cfg = oss.config.load_default()
cfg.credentials_provider = credentials_provider
# Set the region in the configuration.
cfg.region = args.region
# If an endpoint is provided, set the endpoint in the configuration.
if args.endpoint is not None:
cfg.endpoint = args.endpoint
# Create an OSS client based on the configuration.
client = oss.Client(cfg)
# Initiate a multipart upload request to obtain the upload ID for subsequent part uploads.
result = client.initiate_multipart_upload(oss.InitiateMultipartUploadRequest(
bucket=args.bucket,
key=args.key,
))
# Define the size of each part as 5 MB.
part_size = 5 * 1024 * 1024
# Obtain the total size of the file to upload.
data_size = os.path.getsize(args.file_path)
# Initialize the part number, starting from 1.
part_number = 1
# Store the result of each part upload.
upload_parts = []
# Open the file in binary read mode.
with open(args.file_path, 'rb') as f:
# Traverse the file and upload it in parts based on part_size.
for start in range(0, data_size, part_size):
n = part_size
if start + n > data_size: # Handle the case where the last part may be smaller than part_size.
n = data_size - start
# Create a SectionReader to read a specific portion of the file.
reader = oss.io_utils.SectionReader(oss.io_utils.ReadAtReader(f), start, n)
# Upload the part.
up_result = client.upload_part(oss.UploadPartRequest(
bucket=args.bucket,
key=args.key,
upload_id=result.upload_id,
part_number=part_number,
body=reader
))
# Print the result of each part upload.
print(f'status code: {up_result.status_code},'
f' request id: {up_result.request_id},'
f' part number: {part_number},'
f' content md5: {up_result.content_md5},'
f' etag: {up_result.etag},'
f' hash crc64: {up_result.hash_crc64},'
)
# Save the part upload result to the list.
upload_parts.append(oss.UploadPart(part_number=part_number, etag=up_result.etag))
# Increment the part number.
part_number += 1
# Sort the uploaded parts by part number.
parts = sorted(upload_parts, key=lambda p: p.part_number)
# Send a request to complete the multipart upload and merge all parts into a complete object.
result = client.complete_multipart_upload(oss.CompleteMultipartUploadRequest(
bucket=args.bucket,
key=args.key,
upload_id=result.upload_id,
complete_multipart_upload=oss.CompleteMultipartUpload(
parts=parts
)
))
# The following code provides another method to list and merge all part data into a complete object on the server.
# This method is suitable when you are not sure whether all parts are successfully uploaded.
# Merge fragmented data into a complete Object through the server-side List method
# result = client.complete_multipart_upload(oss.CompleteMultipartUploadRequest(
# bucket=args.bucket,
# key=args.key,
# upload_id=result.upload_id,
# complete_all='yes'
# ))
# Print the result of the completed multipart upload.
print(f'status code: {result.status_code},'
f' request id: {result.request_id},'
f' bucket: {result.bucket},'
f' key: {result.key},'
f' location: {result.location},'
f' etag: {result.etag},'
f' encoding type: {result.encoding_type},'
f' hash crc64: {result.hash_crc64},'
f' version id: {result.version_id},'
)
if __name__ == "__main__":
main() # The script entry point. The main function is called when the file is run directly.Common scenarios
References
For the complete sample code for multipart upload, see complete_multipart_upload.py.