You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用AWS Lambda(Node.js)高效复制同账号S3 Bucket全量内容?

Optimizing S3 Bucket Full Replication with AWS Lambda (Node.js)

Great question! Your current approach works for small buckets, but it has some limitations—like not handling pagination (since listObjects only returns up to 1000 objects at a time) and risking early Lambda termination from misplaced callbacks. Let’s break down better ways to replicate an entire bucket.

First, Fix the Pitfalls in Your Current Code

Before jumping to better approaches, let’s address the issues in your existing snippet:

  • Outdated API: listObjects is deprecated—use listObjectsV2 instead, which is the current standard for listing S3 objects.
  • No Pagination Handling: If your bucket has over 1000 objects, your code will only copy the first 1000.
  • Misplaced Callback: Your callback(null, 'All done!') is inside the async.each loop, which will fire before all copies finish, causing Lambda to exit prematurely.
  • Uncontrolled Concurrency: Spamming copyObject calls without rate limiting can trigger S3 API throttling.

Better Approach 1: Optimized SDK v2 Code with Pagination & Controlled Concurrency

This approach uses async/await for cleaner flow, handles pagination to get all objects, and limits concurrent copy operations to avoid throttling.

First, install the p-limit package to control concurrency (run npm i p-limit in your project):

const AWS = require('aws-sdk');
const s3 = new AWS.S3();
const pLimit = require('p-limit');

// Limit concurrent copies to avoid S3 throttling (adjust based on your bucket size)
const concurrencyLimit = 50;
const limit = pLimit(concurrencyLimit);

exports.handler = async (event, context, callback) => {
  const sourceBucket = 'YOUR_SOURCE_BUCKET_NAME';
  const destBucket = 'YOUR_DESTINATION_BUCKET_NAME';
  let allObjects = [];
  let continuationToken = null;

  try {
    // Step 1: Fetch ALL objects from source bucket (handle pagination)
    do {
      const listParams = {
        Bucket: sourceBucket,
        ContinuationToken: continuationToken
      };
      const response = await s3.listObjectsV2(listParams).promise();
      allObjects = [...allObjects, ...response.Contents];
      continuationToken = response.NextContinuationToken;
    } while (continuationToken);

    console.log(`Found ${allObjects.length} objects to replicate`);

    // Step 2: Copy objects with controlled concurrency
    const copyPromises = allObjects.map(object => {
      return limit(async () => {
        const copyParams = {
          CopySource: `${sourceBucket}/${object.Key}`,
          Bucket: destBucket,
          Key: object.Key,
          ACL: 'public-read'
        };
        await s3.copyObject(copyParams).promise();
        console.log(`Successfully copied: ${object.Key}`);
      });
    });

    await Promise.all(copyPromises);
    callback(null, `Replication complete! Copied ${allObjects.length} objects`);
  } catch (err) {
    console.error('Replication failed:', err);
    callback(err);
  }
};

Key Improvements Here:

  • Full Pagination Support: The do-while loop fetches every object in the bucket, even if there are tens of thousands.
  • Controlled Concurrency: p-limit prevents overwhelming S3 with too many simultaneous requests, avoiding throttling.
  • Cleaner Flow: Async/await replaces nested callbacks, making error handling and execution order far easier to follow.
  • Proper Lambda Termination: The final callback only runs after all copies are done, ensuring the function doesn’t exit early.

Better Approach 2: S3 Batch Operations (For Large Buckets)

If your bucket has millions of objects, even optimized Lambda code might hit the 15-minute runtime limit. For this scenario, use S3 Batch Operations:

  1. Generate an inventory list of your source bucket (via S3 Inventory) or create a CSV of object keys.
  2. Use the S3 Control API to create a batch job that runs CopyObject for every object.
  3. You can trigger this job directly from Lambda using the createJob method.

This is far more efficient for massive datasets because S3 handles scaling and concurrency internally, no Lambda runtime limits to worry about.

Important Notes

  • IAM Permissions: Your Lambda execution role needs these permissions:
    • s3:ListBucket on the source bucket
    • s3:GetObject on the source bucket
    • s3:PutObject on the destination bucket
    • (For Batch Operations) s3:CreateJob and s3:GetJobStatus from the S3 Control API
  • Versioning: If you need to replicate object versions, use listObjectVersions instead of listObjectsV2 and add the VersionId parameter to your copy request.
  • SDK v3 Upgrade: For better performance and smaller bundle sizes, consider migrating to AWS SDK v3, which supports modular imports and improved async tools.

内容的提问来源于stack exchange,提问作者Fayza Nawaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 03:46:04