S3技术咨询:不离开S3合并多对象及按桶调整最小分块大小
1. Merging Multiple S3 Objects into One Without Leaving S3
The most efficient way to merge objects entirely within S3 is using the Multipart Upload with UploadPartCopy workflow. Here's how it works:
- Start by initiating a multipart upload for your target object via the
CreateMultipartUploadAPI. You'll get an upload ID to reference for all subsequent steps. - For each source S3 object, use the
UploadPartCopyAPI to copy the object directly as a part of your ongoing multipart upload. This operation keeps data within AWS's internal network, so you don't have to download and re-upload files yourself—saving both time and data transfer costs. - Once all source objects are copied as parts, finalize the upload with the
CompleteMultipartUploadAPI, providing all the part IDs and ETags you received from eachUploadPartCopycall.
Quick tips to keep in mind:
- Always track part IDs and ETags carefully; you'll need them to successfully complete the upload.
- This method avoids cross-region or internet data transfer fees since everything stays within S3.
2. Adjusting Minimum Part Size per Bucket & Handling Small Objects
Can you modify the 5MB minimum part size for a single bucket?
Unfortunately, the short answer is no. The 5MB minimum part size (required for all parts except the final one) is a global S3 rule that can't be adjusted at the bucket, account, or region level. This restriction applies to all multipart uploads, including those using UploadPartCopy.
Workarounds for small objects during merging
If you have source objects smaller than 5MB that aren't the final part of your merge, here are two practical fixes:
- Combine small objects into a temporary ≥5MB object: Use an AWS Lambda function to fetch the small objects, concatenate their content, and upload the result as a temporary S3 object. Then use this temporary object as a valid part in your
UploadPartCopyworkflow. Once the final merged object is ready, you can delete the temporary file to clean up. - Let Lambda handle the full merge: Instead of relying solely on S3 APIs, trigger a Lambda function to read all source objects, merge their content in memory (or use a temporary file for larger totals), and upload the merged result directly as a single object. This works great if the total size of your small objects fits within Lambda's memory and execution time limits.
One exception to note: The final part of a multipart upload can be smaller than 5MB. So if you only have one small object, you can use it as the last part without any extra steps.
内容的提问来源于stack exchange,提问作者Dima Svider

