使用Boto3在EC2上向S3流式传输大文件遇内存问题求助
问题:60GB大文件从HTTP流式传输到S3时内存过载/EC2微实例无法运行
我尝试直接从HTTP服务器将60GB大文件流式传输到S3存储桶,而非先下载再上传。在两个环境中做了测试:
- WSL环境中,内存占用达100%时脚本被终止,即便将
max_concurrency设为2也无效,为何仍会内存过载? - 计划运行代码的EC2(微实例)上,Boto3代码无法运行且无报错,或许需将内存从1GB提至2-3GB?但我希望保留在免费套餐内。
是否有直接流式传输此类大文件的方法?1GB及以下小文件传输完全正常。我认为问题源于内存,代码可能将HTTP文件全读入内存再上传,或许应分块读取并流式传输?但我并非Python专家,已为此研究多日。
以下是我的代码:
def stream_to_s3(self, source_filename, remote_filename): error = 0 self.log(f"====> Streaming {source_filename} to S3://{remote_filename}") s3 = boto3.resource('s3') bucket = s3.Bucket(self.params['UPLOAD_TO_S3']['S3_BUCKET']) destination = bucket.Object(remote_filename) with self.session.get(source_filename, stream=True) as response: GB = 1024 ** 3 MB = 1024 * 1024 max_threshold = 5 * GB # if int(response.headers['content-length']) > max_threshold: TC = TransferConfig(multipart_threshold=max_threshold, max_concurrency=2, multipart_chunksize=8 * MB, use_threads=True) try: destination.upload_fileobj(response.raw, Config=TC) except Exception as e: self.log(f"====> Failure streaming file to S3://{remote_filename}. Reason: {e}") return 1 self.log(f"====> Succeeded streaming file to S3://{remote_filename}")
解决方案
1. 内存过载的核心原因
你的代码存在两个关键问题:
multipart_threshold设为5GB,且注释掉了文件大小判断逻辑,导致小文件触发分块但大文件反而可能被尝试一次性加载;response.raw配合upload_fileobj时,boto3的线程预读机制会持续缓存HTTP响应数据,加上requests默认无缓冲区限制,最终导致内存被占满。
2. 优化后的流式分块上传代码
要实现真正的低内存流式传输,需要手动控制分块读取和上传,避免一次性加载大量数据:
import boto3 from botocore.exceptions import ClientError def stream_to_s3(self, source_filename, remote_filename): self.log(f"====> Streaming {source_filename} to S3://{remote_filename}") s3 = boto3.client('s3') bucket_name = self.params['UPLOAD_TO_S3']['S3_BUCKET'] # 初始化多部分上传 try: mp_upload = s3.create_multipart_upload(Bucket=bucket_name, Key=remote_filename) upload_id = mp_upload['UploadId'] except ClientError as e: self.log(f"====> Failed to initiate multipart upload: {e}") return 1 parts = [] part_number = 1 chunk_size = 8 * 1024 * 1024 # 8MB分块,可根据内存调整为4MB进一步降低占用 success = True try: with self.session.get(source_filename, stream=True) as response: response.raise_for_status() # 逐块读取HTTP响应,严格限制单块内存占用 for chunk in response.iter_content(chunk_size=chunk_size): if not chunk: continue # 上传当前分块到S3 part = s3.upload_part( Bucket=bucket_name, Key=remote_filename, PartNumber=part_number, UploadId=upload_id, Body=chunk ) parts.append({'PartNumber': part_number, 'ETag': part['ETag']}) part_number += 1 self.log(f"====> Uploaded part {part_number-1}") # 完成多部分上传 s3.complete_multipart_upload( Bucket=bucket_name, Key=remote_filename, UploadId=upload_id, MultipartUpload={'Parts': parts} ) self.log(f"====> Succeeded streaming file to S3://{remote_filename}") except Exception as e: self.log(f"====> Failure streaming file to S3://{remote_filename}. Reason: {e}") # 出错时中止上传,避免S3残留无效分块 s3.abort_multipart_upload( Bucket=bucket_name, Key=remote_filename, UploadId=upload_id ) success = False return 0 if success else 1
3. 关键优化点说明
- 使用boto3客户端(
client)而非资源(resource),更灵活控制多部分上传流程; - 用
iter_content(chunk_size=8MB)严格限制每次读取的内存量,每块处理完即释放内存; - 逐块上传S3分块,彻底避免一次性加载大文件到内存;
- 增加错误中止逻辑,防止S3残留未完成的上传分块。
4. EC2微实例适配建议
优化后的代码内存峰值会控制在几十MB以内,完全适配1GB内存的EC2免费套餐实例:
- 可将
chunk_size降至4MB,进一步降低内存占用; - 确保实例网络带宽稳定,避免因网络波动导致分块上传失败。
内容的提问来源于stack exchange,提问作者Yair Glikman
相关产品推荐
相关产品推荐

