如何用Boto或AWS CLI获取S3指定路径下最新5个文件夹及其中文件
获取S3指定路径下最新5个文件夹及文件列表的实现方法
方法一:使用AWS CLI
步骤1:获取最新5个文件夹路径
通过aws s3api结合jq工具,先列出指定前缀下的所有对象,提取文件夹前缀并按最后修改时间排序,筛选出最新的5个:
aws s3api list-objects-v2 --bucket mainbucket --prefix fold1/fold2/fold3/ --query 'Contents[].{Prefix: substring(Key, 0, last_index_of(Key, "/")), LastModified: LastModified}' | jq -r 'group_by(.Prefix) | map({Prefix: .[0].Prefix, LastModified: max_by(.LastModified).LastModified}) | sort_by(.LastModified) | reverse | limit(5; .[]).Prefix'
说明:
list-objects-v2获取对象的键和最后修改时间jq负责分组去重、取每个文件夹的最新修改时间、排序并截取前5个文件夹路径
步骤2:遍历获取每个文件夹的文件列表
将上述结果存入变量,循环遍历每个文件夹并列出文件:
latest_folders=$(aws s3api list-objects-v2 --bucket mainbucket --prefix fold1/fold2/fold3/ --query 'Contents[].{Prefix: substring(Key, 0, last_index_of(Key, "/")), LastModified: LastModified}' | jq -r 'group_by(.Prefix) | map({Prefix: .[0].Prefix, LastModified: max_by(.LastModified).LastModified}) | sort_by(.LastModified) | reverse | limit(5; .[]).Prefix') for folder in $latest_folders; do echo "=== 文件列表:s3://mainbucket/$folder ===" aws s3 ls s3://mainbucket/$folder --recursive done
方法二:使用Python Boto3 SDK
适合需要集成到代码中的场景,以下是完整实现:
import boto3 from collections import defaultdict # 初始化S3客户端 s3 = boto3.client('s3') bucket_name = 'mainbucket' target_prefix = 'fold1/fold2/fold3/' # 收集每个文件夹的最新修改时间 folder_mod_times = defaultdict(lambda: None) paginator = s3.get_paginator('list_objects_v2') # 分页获取所有对象,避免单次请求数据量过大 for page in paginator.paginate(Bucket=bucket_name, Prefix=target_prefix): if 'Contents' not in page: continue for obj in page['Contents']: # 提取文件夹路径(确保以/结尾) folder_path = obj['Key'].rsplit('/', 1)[0] + '/' # 更新文件夹的最新修改时间 current_mod = obj['LastModified'] if folder_mod_times[folder_path] is None or current_mod > folder_mod_times[folder_path]: folder_mod_times[folder_path] = current_mod # 按修改时间倒序排序,取前5个文件夹 sorted_folders = sorted(folder_mod_times.items(), key=lambda x: x[1], reverse=True)[:5] latest_folder_paths = [item[0] for item in sorted_folders] # 遍历每个文件夹,输出文件列表 for folder in latest_folder_paths: print(f"\n=== 文件列表:s3://{bucket_name}/{folder} ===") for page in paginator.paginate(Bucket=bucket_name, Prefix=folder): if 'Contents' not in page: continue for obj in page['Contents']: # 跳过文件夹的占位对象(如果存在) if obj['Key'] == folder: continue print(f" {obj['Key']}")
说明:
- 使用分页器处理大量对象,避免API请求超时
- 用
defaultdict记录每个文件夹的最新修改时间 - 最后遍历筛选出的文件夹,输出所有文件路径
内容的提问来源于stack exchange,提问作者Santhosh
相关产品推荐
相关产品推荐

