如何高效合并S3中的多个Zip文件及小文件至~500MiB量级
Great question—dealing with swarms of tiny S3 files (plus merging existing Zips) without hauling everything to your local machine is a super common pain point, and AWS has several solid, cloud-native ways to pull this off. Let’s break down the most efficient options tailored to your needs:
Lambda is ideal here because it’s serverless—no servers to manage, pay-as-you-go pricing, and it runs entirely in the cloud. Here’s how to structure the workflow:
- Trigger it: Use CloudWatch Events for scheduled batch runs, or S3 Event Notifications if new small files keep arriving.
- Core logic:
- List all small files in your target S3 prefix (filter out the ~500MiB files you want to keep).
- Pull small files into Lambda’s
/tmpdirectory (it supports up to 10GB of temporary storage, more than enough for 500MiB batches). - For regular files: Concatenate content until you hit the ~500MiB threshold, then upload the merged file back to S3 using Multipart Upload (more stable for larger files).
- For Zip files: Use a library like Python’s
zipfileto extract content from small Zips, then repackage into a single large Zip before uploading.
Simplified Python code snippet to get you started:
import boto3 import zipfile from io import BytesIO s3 = boto3.client('s3') BUCKET = 'your-bucket-name' SOURCE_PREFIX = 'raw-files/' TARGET_PREFIX = 'merged-files/' TARGET_SIZE = 500 * 1024 * 1024 # 500MiB def merge_small_files(): # List small files (adjust size threshold as needed) paginator = s3.get_paginator('list_objects_v2') small_files = [] for page in paginator.paginate(Bucket=BUCKET, Prefix=SOURCE_PREFIX): small_files.extend([obj['Key'] for obj in page['Contents'] if obj['Size'] < 1024*1024]) current_batch = [] current_size = 0 batch_num = 1 for file_key in small_files: file_data = s3.get_object(Bucket=BUCKET, Key=file_key)['Body'].read() file_size = len(file_data) # Start a new batch if adding this file exceeds target size if current_size + file_size > TARGET_SIZE and current_batch: merged_key = f"{TARGET_PREFIX}batch_{batch_num}.bin" s3.put_object(Bucket=BUCKET, Key=merged_key, Body=b''.join(current_batch)) current_batch = [] current_size = 0 batch_num += 1 current_batch.append(file_data) current_size += file_size # Upload the final batch if current_batch: merged_key = f"{TARGET_PREFIX}batch_{batch_num}.bin" s3.put_object(Bucket=BUCKET, Key=merged_key, Body=b''.join(current_batch)) def merge_zips(): # List all Zip files paginator = s3.get_paginator('list_objects_v2') zip_files = [] for page in paginator.paginate(Bucket=BUCKET, Prefix=SOURCE_PREFIX): zip_files.extend([obj['Key'] for obj in page['Contents'] if obj['Key'].endswith('.zip')]) merged_zip = BytesIO() with zipfile.ZipFile(merged_zip, 'w') as zf: for zip_key in zip_files: zip_obj = s3.get_object(Bucket=BUCKET, Key=zip_key) with zipfile.ZipFile(BytesIO(zip_obj['Body'].read()), 'r') as source_zf: for file_info in source_zf.infolist(): # Add file to merged Zip (add logic to avoid duplicates if needed) zf.writestr(file_info, source_zf.read(file_info.filename)) # Upload merged Zip to S3 s3.put_object(Bucket=BUCKET, Key=f"{TARGET_PREFIX}merged_zips.zip", Body=merged_zip.getvalue())
Note: Add error handling, pagination, and cleanup (delete/move original files) for production use.
If you’re dealing with hundreds of thousands of small files, Glue’s Spark-based ETL jobs are more efficient than Lambda. Spark is built to handle large-scale file merging:
- Workflow steps:
- Create a Glue Spark job (use Python or Scala).
- Read small files from S3 using Spark’s built-in file readers.
- Use
repartition()orcoalesce()to control output file size (calculate partition count based on total data size to hit ~500MiB per file). - For Zips: Use custom scripts to unpack small Zips, combine their contents, and repackage into large Zips.
Simplified Glue Python snippet:
import sys from awsglue.utils import getResolvedOptions from pyspark.context import SparkContext from awsglue.context import GlueContext from awsglue.job import Job import boto3 import zipfile from io import BytesIO args = getResolvedOptions(sys.argv, ['JOB_NAME']) sc = SparkContext() glueContext = GlueContext(sc) spark = glueContext.spark_session job = Job(glueContext) job.init(args['JOB_NAME'], args) # Merge regular small text files to ~500MiB each df = spark.read.text("s3://your-bucket/raw-files/") # Calculate partitions: total_size / 500MiB (adjust number based on your data) df.repartition(20).write.mode("overwrite").text("s3://your-bucket/merged-files/") # Merge Zip files (simplified for small-scale; use distributed logic for large datasets) s3 = boto3.client('s3') zip_keys = [obj['Key'] for obj in s3.list_objects_v2(Bucket='your-bucket', Prefix='raw-zips/')['Contents']] merged_zip = BytesIO() with zipfile.ZipFile(merged_zip, 'w') as zf: for key in zip_keys: zip_content = s3.get_object(Bucket='your-bucket', Key=key)['Body'].read() with zipfile.ZipFile(BytesIO(zip_content), 'r') as source_zf: for file in source_zf.infolist(): zf.writestr(file, source_zf.read(file.filename)) s3.put_object(Bucket='your-bucket', Key='merged-zips/combined.zip', Body=merged_zip.getvalue()) job.commit()
If you need full control over merging logic (e.g., complex Zip structure rules), spin up an EC2 instance (use spot instances for cost savings):
- Launch an EC2 instance with an EBS volume large enough for temporary storage.
- Sync small files from S3 to EC2:
aws s3 sync s3://your-bucket/raw-files/ /mnt/ebs/temp/ - Use local tools to merge:
- For regular files:
cat /mnt/ebs/temp/small-* > /mnt/ebs/merged/batch1.bin(split into batches to hit 500MiB) - For Zips:
zip -F /mnt/ebs/temp/small-zip*.zip --out /mnt/ebs/merged/combined.zip
- For regular files:
- Sync merged files back to S3:
aws s3 sync /mnt/ebs/merged/ s3://your-bucket/merged-files/ - Terminate the EC2 instance once done to avoid unnecessary costs.
- Filter upfront: Skip files that are already ~500MiB to avoid redundant work.
- Test batch sizing: Adjust batch logic to ensure merged files are as close to 500MiB as possible (account for compression if working with Zips).
- Cleanup strategically: Move original files to an archive prefix or delete them after merging (use S3 lifecycle policies to archive old files long-term).
- Monitor runs: Use CloudWatch Logs to track Lambda/Glue job performance and set alarms for failures.
内容的提问来源于stack exchange,提问作者Pratik Khadloya

