如何修改S3脚本:获取最新创建目录并下载其内所有文件
问题描述
我在AWS S3中有路径 object1/object2/object3/object4/,该路径下包含多个子目录,示例如下:
directory1/directory2/directory3/directory4/2022-30-09-15h21/ directory1/directory2/directory3/directory4/2023-20-12-12h30/ directory1/directory2/directory3/directory4/2022-31-12-09h34/ directory1/directory2/directory3/directory4/2023-12-08-14h56/
我希望选中directory4/下最新创建的目录,并下载该目录内的所有文件。当前编写的Python脚本仅能获取该最新目录下最新创建的单个文件,需要修改脚本实现需求。原脚本如下:
import boto3 from datetime import datetime session_root = boto3.Session(region_name='eu-west-3', profile_name='my_profile') s3_client = session_root.client('s3') bucket_name = 'my_bucket' prefix = 'object1/object2/object3/object4/' # List objects in the bucket response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=prefix) # Extract the object names and convert them to datetime objects objects_with_dates = [(obj['Key'], datetime.strptime(obj['LastModified'].strftime('%Y-%m-%d %H:%M:%S'), '%Y-%m-%d %H:%M:%S')) for obj in response.get('Contents', [])] # Find the latest created object latest_object = max(objects_with_dates, key=lambda x: x[1]) print("Last created S3 object:", latest_object[0]) # 返回结果示例:object1/object2/object3/object4/2023-20-12-12h30/my_file.csv
解决方案
要实现需求,需要先定位到directory4/下的最新日期目录,再批量下载该目录下的所有文件,修改后的完整脚本如下:
import boto3 from datetime import datetime import os # 初始化S3客户端 session_root = boto3.Session(region_name='eu-west-3', profile_name='my_profile') s3_client = session_root.client('s3') bucket_name = 'my_bucket' prefix = 'object1/object2/object3/object4/' # 1. 遍历所有对象,按目录分组并记录每个目录的最新修改时间 directory_time_map = {} response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=prefix) while True: contents = response.get('Contents', []) for obj in contents: # 跳过前缀本身的空对象 if obj['Key'] == prefix: continue # 提取对象所属的日期目录 relative_path = obj['Key'][len(prefix):] directory_name = prefix + relative_path.split('/')[0] + '/' # 更新目录的最新修改时间 last_modified = obj['LastModified'] if directory_name not in directory_time_map or last_modified > directory_time_map[directory_name]: directory_time_map[directory_name] = last_modified # 处理分页(对象数量超过1000时触发) if not response.get('IsTruncated'): break response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=prefix, ContinuationToken=response['NextContinuationToken']) # 2. 确定最新的目标目录 if not directory_time_map: print("未找到任何子目录") exit() latest_directory = max(directory_time_map, key=directory_time_map.get) print("最新目录:", latest_directory) # 3. 批量下载该目录下的所有文件 response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=latest_directory) while True: contents = response.get('Contents', []) for obj in contents: # 跳过S3模拟目录的空对象 if obj['Size'] == 0: continue # 构建本地保存路径,保持与S3一致的目录结构 local_path = obj['Key'].replace(prefix, '') local_dir = os.path.dirname(local_path) # 创建本地目录(不存在则自动创建) if not os.path.exists(local_dir): os.makedirs(local_dir, exist_ok=True) # 下载文件 s3_client.download_file(bucket_name, obj['Key'], local_path) print(f"已下载: {obj['Key']} -> {local_path}") # 处理分页 if not response.get('IsTruncated'): break response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=latest_directory, ContinuationToken=response['NextContinuationToken']) print("所有文件下载完成")
关键修改说明
- 分页处理:添加分页逻辑,避免因S3返回对象数量超过1000条导致的内容遗漏。
- 目录分组:通过分割对象Key提取所属目录,确保每个目录的最新时间是其下所有文件的最晚修改时间。
- 本地目录同步:自动创建与S3对应的本地目录结构,避免文件保存路径错误。
- 过滤空对象:跳过S3中模拟目录的空对象,只下载实际文件。
内容的提问来源于stack exchange,提问作者mlflow
相关产品推荐
相关产品推荐

