如何用AWS Lambda(Node.js)高效复制同账号S3 Bucket全量内容?
Great question! Your current approach works for small buckets, but it has some limitations—like not handling pagination (since listObjects only returns up to 1000 objects at a time) and risking early Lambda termination from misplaced callbacks. Let’s break down better ways to replicate an entire bucket.
First, Fix the Pitfalls in Your Current Code
Before jumping to better approaches, let’s address the issues in your existing snippet:
- Outdated API:
listObjectsis deprecated—uselistObjectsV2instead, which is the current standard for listing S3 objects. - No Pagination Handling: If your bucket has over 1000 objects, your code will only copy the first 1000.
- Misplaced Callback: Your
callback(null, 'All done!')is inside theasync.eachloop, which will fire before all copies finish, causing Lambda to exit prematurely. - Uncontrolled Concurrency: Spamming
copyObjectcalls without rate limiting can trigger S3 API throttling.
Better Approach 1: Optimized SDK v2 Code with Pagination & Controlled Concurrency
This approach uses async/await for cleaner flow, handles pagination to get all objects, and limits concurrent copy operations to avoid throttling.
First, install the p-limit package to control concurrency (run npm i p-limit in your project):
const AWS = require('aws-sdk'); const s3 = new AWS.S3(); const pLimit = require('p-limit'); // Limit concurrent copies to avoid S3 throttling (adjust based on your bucket size) const concurrencyLimit = 50; const limit = pLimit(concurrencyLimit); exports.handler = async (event, context, callback) => { const sourceBucket = 'YOUR_SOURCE_BUCKET_NAME'; const destBucket = 'YOUR_DESTINATION_BUCKET_NAME'; let allObjects = []; let continuationToken = null; try { // Step 1: Fetch ALL objects from source bucket (handle pagination) do { const listParams = { Bucket: sourceBucket, ContinuationToken: continuationToken }; const response = await s3.listObjectsV2(listParams).promise(); allObjects = [...allObjects, ...response.Contents]; continuationToken = response.NextContinuationToken; } while (continuationToken); console.log(`Found ${allObjects.length} objects to replicate`); // Step 2: Copy objects with controlled concurrency const copyPromises = allObjects.map(object => { return limit(async () => { const copyParams = { CopySource: `${sourceBucket}/${object.Key}`, Bucket: destBucket, Key: object.Key, ACL: 'public-read' }; await s3.copyObject(copyParams).promise(); console.log(`Successfully copied: ${object.Key}`); }); }); await Promise.all(copyPromises); callback(null, `Replication complete! Copied ${allObjects.length} objects`); } catch (err) { console.error('Replication failed:', err); callback(err); } };
Key Improvements Here:
- Full Pagination Support: The
do-whileloop fetches every object in the bucket, even if there are tens of thousands. - Controlled Concurrency:
p-limitprevents overwhelming S3 with too many simultaneous requests, avoiding throttling. - Cleaner Flow: Async/await replaces nested callbacks, making error handling and execution order far easier to follow.
- Proper Lambda Termination: The final callback only runs after all copies are done, ensuring the function doesn’t exit early.
Better Approach 2: S3 Batch Operations (For Large Buckets)
If your bucket has millions of objects, even optimized Lambda code might hit the 15-minute runtime limit. For this scenario, use S3 Batch Operations:
- Generate an inventory list of your source bucket (via S3 Inventory) or create a CSV of object keys.
- Use the S3 Control API to create a batch job that runs
CopyObjectfor every object. - You can trigger this job directly from Lambda using the
createJobmethod.
This is far more efficient for massive datasets because S3 handles scaling and concurrency internally, no Lambda runtime limits to worry about.
Important Notes
- IAM Permissions: Your Lambda execution role needs these permissions:
s3:ListBucketon the source buckets3:GetObjecton the source buckets3:PutObjecton the destination bucket- (For Batch Operations)
s3:CreateJobands3:GetJobStatusfrom the S3 Control API
- Versioning: If you need to replicate object versions, use
listObjectVersionsinstead oflistObjectsV2and add theVersionIdparameter to your copy request. - SDK v3 Upgrade: For better performance and smaller bundle sizes, consider migrating to AWS SDK v3, which supports modular imports and improved async tools.
内容的提问来源于stack exchange,提问作者Fayza Nawaz

