如何通过单次请求读取S3存储桶中10万张图片以减少API调用?
问题:S3批量读取10万张图片的IO/API请求优化
我有一个存储了10万张PNG/JPG格式图片的S3存储桶,希望通过单次S3请求读取所有图片。以下是我的代码:
# geting connection to s3 s3 = boto3.resource('s3') bucket = s3.Bucket(bucket_name) # geting list of all objectsummery in the relevant dir - this is a request to the api crops_path_list = bucket.objects.filter(Prefix=prefix) # iterating on the list and geting objects for index, crop in enumerate(crops_path_list): # getting the bytes from s3 - this a request to the api crop = crop.get() # working with the bytes to make them image image_content = crop['Body'].read() bytes_images = BytesIO(image_content) image = Image.open(bytes_images) image = image.convert("RGB") image = np.asarray(image) image = np.ascontiguousarray(image.transpose(2, 0, 1)) image = torch.from_numpy(np.array(image)).unsqueeze(0).to(dtype=torch.float32, device=device) # adding to list of all images images.append(image)
目前这段代码耗时极长,因为每次调用.get()都会发起一次API请求。我未找到能通过单次请求读取整个crops_path_list的方法,除了使用线程、子进程或压缩文件的方式,请问还有其他减少IO和API请求的方案吗?
可行优化方案
1. 切换到boto3 Client复用连接池
使用boto3.client而非resource,Client底层默认维护HTTP连接池,可复用TCP连接,减少每次请求的握手开销,间接降低IO耗时。示例代码:
import boto3 from io import BytesIO from PIL import Image import numpy as np import torch s3_client = boto3.client('s3') paginator = s3_client.get_paginator('list_objects_v2') page_iterator = paginator.paginate(Bucket=bucket_name, Prefix=prefix) images = [] device = torch.device("cuda" if torch.cuda.is_available() else "cpu") for page in page_iterator: if 'Contents' not in page: continue for obj in page['Contents']: response = s3_client.get_object(Bucket=bucket_name, Key=obj['Key']) image_content = response['Body'].read() bytes_images = BytesIO(image_content) image = Image.open(bytes_images).convert("RGB") image = np.asarray(image) image = np.ascontiguousarray(image.transpose(2, 0, 1)) image = torch.from_numpy(image).unsqueeze(0).to(dtype=torch.float32, device=device) images.append(image)
2. 启用S3 Transfer Acceleration
若客户端与S3桶不在同一AWS区域,开启Transfer Acceleration可利用AWS全球边缘网络加速数据传输,降低跨区域延迟。只需修改客户端配置:
s3_client = boto3.client('s3', config=boto3.client.Config(use_accelerate_endpoint=True))
3. 优化图片预处理流程
减少中间转换步骤,降低CPU/内存瓶颈,避免拖慢IO处理速度:
- 使用
torchvision.io.read_image直接读取图片为张量,省略PIL转RGB、numpy数组转换等步骤:
from torchvision.io import read_image # 替换原图片处理逻辑 bytes_images = BytesIO(image_content) image = read_image(bytes_images).unsqueeze(0).to(dtype=torch.float32, device=device) # read_image默认返回CHW格式张量,无需手动transpose和类型转换
- 延迟设备转移:先在CPU上批量处理所有图片,再一次性转移到GPU,减少多次设备同步开销。
4. 使用S3 Inventory批量获取对象列表
对于10万级别的对象,使用S3 Inventory预先导出桶内对象的元数据到CSV/Parquet文件,直接读取该文件获取对象Key,避免多次调用ListObjects API。导出后只需下载Inventory文件解析即可,大幅减少列表请求次数。
内容的提问来源于stack exchange,提问作者Benny Semyonov
相关产品推荐
相关产品推荐

