You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过单次请求读取S3存储桶中10万张图片以减少API调用?

问题:S3批量读取10万张图片的IO/API请求优化

我有一个存储了10万张PNG/JPG格式图片的S3存储桶,希望通过单次S3请求读取所有图片。以下是我的代码:

# geting connection to s3
s3 = boto3.resource('s3')
bucket = s3.Bucket(bucket_name)

# geting list of all objectsummery in the relevant dir - this is a request to the api
crops_path_list = bucket.objects.filter(Prefix=prefix)

# iterating on the list and geting objects
for index, crop in enumerate(crops_path_list):
    # getting the bytes from s3 - this a request to the api
    crop = crop.get()

    # working with the bytes to make them image
    image_content = crop['Body'].read()
    bytes_images = BytesIO(image_content)
    image = Image.open(bytes_images)
    image = image.convert("RGB")
    image = np.asarray(image)
    image = np.ascontiguousarray(image.transpose(2, 0, 1))
    image  = torch.from_numpy(np.array(image)).unsqueeze(0).to(dtype=torch.float32, device=device)

    # adding to list of all images
    images.append(image)

目前这段代码耗时极长,因为每次调用.get()都会发起一次API请求。我未找到能通过单次请求读取整个crops_path_list的方法,除了使用线程、子进程或压缩文件的方式,请问还有其他减少IO和API请求的方案吗?

可行优化方案

1. 切换到boto3 Client复用连接池

使用boto3.client而非resource,Client底层默认维护HTTP连接池,可复用TCP连接,减少每次请求的握手开销,间接降低IO耗时。示例代码:

import boto3
from io import BytesIO
from PIL import Image
import numpy as np
import torch

s3_client = boto3.client('s3')
paginator = s3_client.get_paginator('list_objects_v2')
page_iterator = paginator.paginate(Bucket=bucket_name, Prefix=prefix)

images = []
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")

for page in page_iterator:
    if 'Contents' not in page:
        continue
    for obj in page['Contents']:
        response = s3_client.get_object(Bucket=bucket_name, Key=obj['Key'])
        image_content = response['Body'].read()
        bytes_images = BytesIO(image_content)
        image = Image.open(bytes_images).convert("RGB")
        image = np.asarray(image)
        image = np.ascontiguousarray(image.transpose(2, 0, 1))
        image = torch.from_numpy(image).unsqueeze(0).to(dtype=torch.float32, device=device)
        images.append(image)

2. 启用S3 Transfer Acceleration

若客户端与S3桶不在同一AWS区域,开启Transfer Acceleration可利用AWS全球边缘网络加速数据传输,降低跨区域延迟。只需修改客户端配置:

s3_client = boto3.client('s3', config=boto3.client.Config(use_accelerate_endpoint=True))

3. 优化图片预处理流程

减少中间转换步骤,降低CPU/内存瓶颈,避免拖慢IO处理速度:

  • 使用torchvision.io.read_image直接读取图片为张量,省略PIL转RGB、numpy数组转换等步骤:
from torchvision.io import read_image

# 替换原图片处理逻辑
bytes_images = BytesIO(image_content)
image = read_image(bytes_images).unsqueeze(0).to(dtype=torch.float32, device=device)
# read_image默认返回CHW格式张量,无需手动transpose和类型转换
  • 延迟设备转移:先在CPU上批量处理所有图片,再一次性转移到GPU,减少多次设备同步开销。

4. 使用S3 Inventory批量获取对象列表

对于10万级别的对象,使用S3 Inventory预先导出桶内对象的元数据到CSV/Parquet文件,直接读取该文件获取对象Key,避免多次调用ListObjects API。导出后只需下载Inventory文件解析即可,大幅减少列表请求次数。

内容的提问来源于stack exchange,提问作者Benny Semyonov

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.01 22:41:13