如何无需临时文件,通过requests+boto3将文件直接上传至AWS S3?
可以跳过磁盘写入,直接流式上传到S3
当然能优化,boto3完全支持直接从网络流上传到S3,不用先写磁盘,尤其是大文件场景,流式处理能大幅降低内存占用和IO开销。
两种实现方案
1. 小文件:直接用put_object
如果文件体积不大,能轻松放进内存,直接把requests返回的内容传给put_object即可:
import requests import boto3 s3_client = boto3.client('s3') bucket_name = "your-bucket-name" s3_key = "path/to/save/file.ext" # 下载文件到内存 response = requests.get("https://example.com/your-file.ext") response.raise_for_status() # 确保请求成功 # 直接上传到S3 s3_client.put_object(Bucket=bucket_name, Key=s3_key, Body=response.content)
2. 大文件:流式上传(重点推荐)
针对大文件,上面的方法会把整个文件加载到内存,容易出现内存溢出。这时候用requests的流式请求+boto3的upload_fileobj方法,分块读取上传,全程不会占用过多内存:
import requests import boto3 s3_client = boto3.client('s3') bucket_name = "your-bucket-name" s3_key = "path/to/save/file.ext" # 开启流式请求,不一次性加载整个文件到内存 with requests.get("https://example.com/large-file.ext", stream=True) as response: response.raise_for_status() # 直接把响应流传给S3的upload_fileobj s3_client.upload_fileobj(response, bucket_name, s3_key)
为什么这能行?
requests.get(stream=True)会让响应以流的形式返回,每次只读取一小部分数据,不会把整个文件塞进内存。upload_fileobj方法接受任何实现了read()方法的类文件对象,而requests的Response对象正好满足这个要求——它会自动分块读取并上传到S3,底层还会自动处理大文件的分块上传逻辑(比如超过100MB的文件会自动用multipart upload)。
额外优化点
- 可以给
upload_fileobj加Config参数,自定义分块大小和并发数,比如:from boto3.s3.transfer import TransferConfig config = TransferConfig(multipart_threshold=1024*1024*20, # 20MB以上自动分块 max_concurrency=10, multipart_chunksize=1024*1024*5) # 每个分块5MB s3_client.upload_fileobj(response, bucket_name, s3_key, Config=config) - 建议加上异常捕获,处理网络中断、S3上传失败等情况,避免中途出错没有反馈。
内容的提问来源于stack exchange,提问作者user2138149
相关产品推荐
相关产品推荐

