You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用boto3高效上传大量文件至Amazon S3?

当然可以!用多线程/多进程批量上传S3,速度直接起飞

上传上万份10MB的文件,串行方式确实会把大量时间耗在等待网络响应上,并行上传正好能把这些“等待时间”利用起来,大幅提升效率。下面给你分享几种实战方案,附代码示例,你可以直接套用:

先搞懂:为什么串行上传慢?

每次上传单个文件时,都要经历建立连接→发送请求→等待S3响应→关闭连接的流程,串行模式下这些步骤只能挨个执行,大部分时间都在等网络,CPU和带宽都没充分利用。并行模式下,多个文件的上传流程同时进行,能把带宽和资源拉满。

方案1:多线程上传(最推荐,适合纯IO密集型任务)

上传文件属于网络IO密集型任务,Python的多线程(不受GIL限制)比多进程更高效(进程切换开销大)。用concurrent.futures.ThreadPoolExecutor可以轻松实现批量并行上传:

代码示例

import boto3
from concurrent.futures import ThreadPoolExecutor, as_completed
import os
from botocore.config import Config

# 初始化带重试配置的S3客户端(避免网络波动导致失败)
s3_config = Config(
    retries={
        'max_attempts': 5,  # 最多重试5次
        'mode': 'standard'  # 标准重试模式
    }
)
s3 = boto3.client('s3', config=s3_config)

# 配置你的参数
BUCKET_NAME = "your-target-bucket"
LOCAL_FILE_DIR = "/path/to/your/local/files"

def upload_single_file(local_file_path):
    """上传单个文件到S3,返回上传结果"""
    try:
        # 用本地文件名作为S3对象键,你也可以自定义路径
        s3_object_key = os.path.basename(local_file_path)
        s3.upload_file(local_file_path, BUCKET_NAME, s3_object_key)
        print(f"✅ 成功上传: {s3_object_key}")
        return True
    except Exception as e:
        print(f"❌ 上传失败 {local_file_path}: {str(e)}")
        return False

def batch_upload():
    # 获取本地目录下所有文件(过滤掉文件夹)
    all_files = [
        os.path.join(LOCAL_FILE_DIR, filename)
        for filename in os.listdir(LOCAL_FILE_DIR)
        if os.path.isfile(os.path.join(LOCAL_FILE_DIR, filename))
    ]
    total_files = len(all_files)
    print(f"待上传文件总数: {total_files}")

    # 配置线程池大小,建议先测试50-100,太大可能触发S3限流
    with ThreadPoolExecutor(max_workers=80) as executor:
        # 批量提交上传任务
        future_to_file = {executor.submit(upload_single_file, file): file for file in all_files}
        
        # 跟踪上传进度
        completed_count = 0
        for future in as_completed(future_to_file):
            completed_count += 1
            print(f"进度: {completed_count}/{total_files}")

if __name__ == "__main__":
    batch_upload()

注意事项

  • 线程数不要盲目调大:S3有请求速率限制(默认每个前缀每秒3500个PUT/POST请求),如果线程数太大,可能会触发限流(返回429错误),建议从50开始测试,逐步调整。
  • 加重试机制:网络波动很常见,给boto3客户端配置重试能避免少量失败导致整个任务中断。

方案2:多进程上传(适合带CPU密集型预处理的场景)

如果你的文件上传前需要压缩、加密、格式转换等CPU密集型操作,那多进程更合适(避开Python GIL的限制)。只需要把上面代码里的ThreadPoolExecutor换成ProcessPoolExecutor即可:

代码示例(核心部分改动)

from concurrent.futures import ProcessPoolExecutor  # 替换ThreadPoolExecutor

# ...其他代码不变...

with ProcessPoolExecutor(max_workers=8) as executor:  # 进程数建议和CPU核心数相当
    future_to_file = {executor.submit(upload_single_file, file): file for file in all_files}
    # ...进度跟踪部分不变...

注意事项

  • 进程数建议设置为CPU核心数的1-2倍,过多进程会导致CPU上下文切换开销增大。
  • Windows系统下,必须把主逻辑放在if __name__ == "__main__":块里,否则会出现进程启动异常。

额外优化技巧

  1. 使用S3 Transfer Config优化单文件上传:对于10MB的文件,可以开启分段上传进一步提速,给upload_file加配置:
    from boto3.s3.transfer import TransferConfig
    
    transfer_config = TransferConfig(
        multipart_threshold=8 * 1024 * 1024,  # 大于8MB自动分段
        max_concurrency=10,  # 单文件分段上传的并发数
        multipart_chunksize=8 * 1024 * 1024  # 每个分段大小8MB
    )
    s3.upload_file(local_file_path, BUCKET_NAME, s3_object_key, Config=transfer_config)
    
  2. 批量命名S3对象:如果需要按目录结构上传,不要直接用basename,而是保留相对路径作为S3对象键,比如os.path.relpath(local_file_path, LOCAL_FILE_DIR)。
  3. 失败文件重传:可以把上传失败的文件记录到日志,最后单独处理,避免遗漏。

内容的提问来源于stack exchange,提问作者Shivaraj Nesargi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:19:22