如何使用AWS Lambda高效解压S3中的大型Zip文件?
解决Lambda处理S3大Zip文件内存溢出问题
原代码的核心问题是将整个Zip文件一次性加载到内存,当文件体积超过Lambda的内存配额时,直接触发MemoryError。以下是无需加载整个文件的分块处理方案:
核心思路
利用Zip文件的结构特性,仅读取必要的元数据(中央目录),然后针对每个文件单独从S3读取对应的压缩数据块,解压后直接上传到目标桶,全程不加载完整Zip文件到内存。
修改后的代码
import boto3 import zipfile from io import BytesIO def lambda_handler(event, context): s3 = boto3.client('s3') source_bucket = 'bucket1' target_bucket = 'bucket2' # 分页遍历源桶对象(如果是S3事件触发,建议直接从event提取对象Key) paginator = s3.get_paginator('list_objects_v2') for page in paginator.paginate(Bucket=source_bucket): for obj in page.get('Contents', []): key = obj['Key'] if not key.endswith('.zip'): print(f'{key} is not a zip file.') continue file_size = obj['Size'] if file_size == 0: print(f'{key} is an empty zip file.') continue # 1. 读取Zip文件末尾,获取中央目录位置 end_range = min(file_size, 65536) end_response = s3.get_object( Bucket=source_bucket, Key=key, Range=f'bytes={file_size - end_range}-' ) end_data = end_response['Body'].read() eocd = zipfile._EndRecData(end_data) cd_start = eocd[zipfile._ECD_CD_START] cd_size = eocd[zipfile._ECD_CD_SIZE] # 2. 读取中央目录,解析所有文件元数据 cd_response = s3.get_object( Bucket=source_bucket, Key=key, Range=f'bytes={cd_start}-{cd_start + cd_size - 1}' ) cd_data = cd_response['Body'].read() cd = zipfile._CentralDirectory(cd_data, 0, len(cd_data), cd_start, file_size) # 3. 逐个处理Zip内的文件 for zinfo in cd._parse(): if zinfo.is_dir(): continue # 计算当前文件压缩数据在S3中的字节范围 header_len = len(zinfo.FileHeader()) data_start = zinfo.header_offset + header_len data_end = data_start + zinfo.compress_size - 1 # 读取单个文件的压缩数据块 file_response = s3.get_object( Bucket=source_bucket, Key=key, Range=f'bytes={data_start}-{data_end}' ) compressed_data = file_response['Body'].read() # 解压并上传到目标桶 if zinfo.compress_type == zipfile.ZIP_STORED: # 未压缩文件直接上传 s3.put_object( Bucket=target_bucket, Key=zinfo.filename, Body=compressed_data ) else: # 压缩文件先解压再上传 with BytesIO(compressed_data) as comp_stream: with zipfile.ZipFile(comp_stream, 'r') as zf: with zf.open(zinfo.filename) as uncomp_stream: s3.upload_fileobj(uncomp_stream, target_bucket, zinfo.filename) print(f'Successfully extracted {key} to {target_bucket}')
关键优化点
- 避免全量加载:仅读取Zip的元数据和单个文件的压缩块,内存占用仅取决于单个文件的压缩大小,而非整个Zip文件体积。
- S3 Range请求:通过指定字节范围,精准读取所需数据,减少不必要的网络传输和内存消耗。
- 分文件处理:逐个处理Zip内的文件,即使包含大量小文件,也能控制内存占用。
注意事项
- 如果Lambda由S3事件触发,建议直接从
event参数中提取触发的对象Key,无需遍历整个桶,提升执行效率。 - 注意Lambda的15分钟执行时间限制,若Zip包含超大量文件,可结合Step Functions分批处理。
- 可添加重试逻辑处理S3临时请求错误,例如使用
tenacity库(需部署Lambda层)。
内容的提问来源于stack exchange,提问作者Ali AzG
相关产品推荐
相关产品推荐

