如何使用aioboto3异步快速获取Amazon S3指定深度的最底层子文件夹
优化方案说明
核心优化点是利用S3 API原生的Delimiter参数,不需要拉取全量对象,直接由S3服务端返回层级前缀(也就是你要的文件夹),请求返回的数据量会比拉全量Contents少非常多,速度提升效果极其明显。
具体实现逻辑
- 你要获取深度为N的路径,直接逐层遍历前缀即可,不需要处理任何文件对象。比如你要深度为4的文件夹,只需要遍历4层前缀,每一层都用
Delimiter='/'查询,返回的CommonPrefixes字段就是当前层级的子文件夹,完全不用拉取文件的Contents列表。 - 相比你现在拉全量对象再过滤的方案,请求次数和数据传输量都会降低几个数量级,尤其当你的桶内文件数量远多于文件夹时,优化效果更突出。
适配异步场景的代码示例
修改后的实现不需要原来的get_paths_by_depth函数,参考代码如下:
from typing import Set from collections import deque async def get_subfolders_by_depth(self, bucket: str, root_prefix: str, target_depth: int) -> Set[str]: subfolders = set() # 计算根前缀的初始深度,可根据自身的前缀规则调整 if root_prefix == "": current_root_depth = 0 else: current_root_depth = root_prefix.count("/") if root_prefix.endswith("/") else root_prefix.count("/") + 1 queue = deque([(root_prefix, current_root_depth)]) while queue: current_prefix, current_depth = queue.popleft() # 到达目标深度直接收集结果 if current_depth == target_depth: subfolders.add(current_prefix.rstrip("/")) continue # 超过目标深度直接跳过 if current_depth > target_depth: continue # 携带Delimiter查询,仅返回前缀,不返回文件对象 result = await self.s3_client.list_objects_v2( Bucket=bucket, Prefix=current_prefix, Delimiter="/" ) # 处理当前页的子前缀 if "CommonPrefixes" in result: for prefix_item in result["CommonPrefixes"]: child_prefix = prefix_item["Prefix"] queue.append((child_prefix, current_depth + 1)) # 分页处理 while result["IsTruncated"]: result = await self.s3_client.list_objects_v2( Bucket=bucket, Prefix=current_prefix, Delimiter="/", ContinuationToken=result["NextContinuationToken"] ) if "CommonPrefixes" in result: for prefix_item in result["CommonPrefixes"]: child_prefix = prefix_item["Prefix"] queue.append((child_prefix, current_depth + 1)) return subfolders
使用说明
你原来的外层调用逻辑不需要大改,把原来的get_subfolders替换为上述实现,传入目标深度(比如你之前使用的depth=4)即可。整个查询过程不会拉取任何文件的元数据,所有返回结果都是S3服务端整理好的文件夹前缀,数据量极小,分页次数也会大幅降低。
内容的提问来源于stack exchange,提问作者KZiovas
相关产品推荐
相关产品推荐

