如何使用boto3高效上传大量文件至Amazon S3?
当然可以!用多线程/多进程批量上传S3,速度直接起飞
上传上万份10MB的文件,串行方式确实会把大量时间耗在等待网络响应上,并行上传正好能把这些“等待时间”利用起来,大幅提升效率。下面给你分享几种实战方案,附代码示例,你可以直接套用:
先搞懂:为什么串行上传慢?
每次上传单个文件时,都要经历建立连接→发送请求→等待S3响应→关闭连接的流程,串行模式下这些步骤只能挨个执行,大部分时间都在等网络,CPU和带宽都没充分利用。并行模式下,多个文件的上传流程同时进行,能把带宽和资源拉满。
方案1:多线程上传(最推荐,适合纯IO密集型任务)
上传文件属于网络IO密集型任务,Python的多线程(不受GIL限制)比多进程更高效(进程切换开销大)。用concurrent.futures.ThreadPoolExecutor可以轻松实现批量并行上传:
代码示例
import boto3 from concurrent.futures import ThreadPoolExecutor, as_completed import os from botocore.config import Config # 初始化带重试配置的S3客户端(避免网络波动导致失败) s3_config = Config( retries={ 'max_attempts': 5, # 最多重试5次 'mode': 'standard' # 标准重试模式 } ) s3 = boto3.client('s3', config=s3_config) # 配置你的参数 BUCKET_NAME = "your-target-bucket" LOCAL_FILE_DIR = "/path/to/your/local/files" def upload_single_file(local_file_path): """上传单个文件到S3,返回上传结果""" try: # 用本地文件名作为S3对象键,你也可以自定义路径 s3_object_key = os.path.basename(local_file_path) s3.upload_file(local_file_path, BUCKET_NAME, s3_object_key) print(f"✅ 成功上传: {s3_object_key}") return True except Exception as e: print(f"❌ 上传失败 {local_file_path}: {str(e)}") return False def batch_upload(): # 获取本地目录下所有文件(过滤掉文件夹) all_files = [ os.path.join(LOCAL_FILE_DIR, filename) for filename in os.listdir(LOCAL_FILE_DIR) if os.path.isfile(os.path.join(LOCAL_FILE_DIR, filename)) ] total_files = len(all_files) print(f"待上传文件总数: {total_files}") # 配置线程池大小,建议先测试50-100,太大可能触发S3限流 with ThreadPoolExecutor(max_workers=80) as executor: # 批量提交上传任务 future_to_file = {executor.submit(upload_single_file, file): file for file in all_files} # 跟踪上传进度 completed_count = 0 for future in as_completed(future_to_file): completed_count += 1 print(f"进度: {completed_count}/{total_files}") if __name__ == "__main__": batch_upload()
注意事项
- 线程数不要盲目调大:S3有请求速率限制(默认每个前缀每秒3500个PUT/POST请求),如果线程数太大,可能会触发限流(返回429错误),建议从50开始测试,逐步调整。
- 加重试机制:网络波动很常见,给boto3客户端配置重试能避免少量失败导致整个任务中断。
方案2:多进程上传(适合带CPU密集型预处理的场景)
如果你的文件上传前需要压缩、加密、格式转换等CPU密集型操作,那多进程更合适(避开Python GIL的限制)。只需要把上面代码里的ThreadPoolExecutor换成ProcessPoolExecutor即可:
代码示例(核心部分改动)
from concurrent.futures import ProcessPoolExecutor # 替换ThreadPoolExecutor # ...其他代码不变... with ProcessPoolExecutor(max_workers=8) as executor: # 进程数建议和CPU核心数相当 future_to_file = {executor.submit(upload_single_file, file): file for file in all_files} # ...进度跟踪部分不变...
注意事项
- 进程数建议设置为CPU核心数的1-2倍,过多进程会导致CPU上下文切换开销增大。
- Windows系统下,必须把主逻辑放在
if __name__ == "__main__":块里,否则会出现进程启动异常。
额外优化技巧
- 使用S3 Transfer Config优化单文件上传:对于10MB的文件,可以开启分段上传进一步提速,给
upload_file加配置:from boto3.s3.transfer import TransferConfig transfer_config = TransferConfig( multipart_threshold=8 * 1024 * 1024, # 大于8MB自动分段 max_concurrency=10, # 单文件分段上传的并发数 multipart_chunksize=8 * 1024 * 1024 # 每个分段大小8MB ) s3.upload_file(local_file_path, BUCKET_NAME, s3_object_key, Config=transfer_config) - 批量命名S3对象:如果需要按目录结构上传,不要直接用basename,而是保留相对路径作为S3对象键,比如
os.path.relpath(local_file_path, LOCAL_FILE_DIR)。 - 失败文件重传:可以把上传失败的文件记录到日志,最后单独处理,避免遗漏。
内容的提问来源于stack exchange,提问作者Shivaraj Nesargi
相关产品推荐
相关产品推荐

