You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修改S3脚本:获取最新创建目录并下载其内所有文件

问题描述

我在AWS S3中有路径 object1/object2/object3/object4/,该路径下包含多个子目录,示例如下:

directory1/directory2/directory3/directory4/2022-30-09-15h21/
directory1/directory2/directory3/directory4/2023-20-12-12h30/
directory1/directory2/directory3/directory4/2022-31-12-09h34/
directory1/directory2/directory3/directory4/2023-12-08-14h56/

我希望选中directory4/下最新创建的目录,并下载该目录内的所有文件。当前编写的Python脚本仅能获取该最新目录下最新创建的单个文件,需要修改脚本实现需求。原脚本如下:

import boto3
from datetime import datetime 

session_root = boto3.Session(region_name='eu-west-3', profile_name='my_profile')
s3_client = session_root.client('s3') 

bucket_name = 'my_bucket' 

prefix = 'object1/object2/object3/object4/'

# List objects in the bucket 
response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=prefix) 

# Extract the object names and convert them to datetime objects 

objects_with_dates = [(obj['Key'], datetime.strptime(obj['LastModified'].strftime('%Y-%m-%d %H:%M:%S'), '%Y-%m-%d %H:%M:%S')) for obj in response.get('Contents', [])] 
# Find the latest created object 

latest_object = max(objects_with_dates, key=lambda x: x[1]) 

print("Last created S3 object:", latest_object[0]) # 返回结果示例:object1/object2/object3/object4/2023-20-12-12h30/my_file.csv
解决方案

要实现需求,需要先定位到directory4/下的最新日期目录,再批量下载该目录下的所有文件,修改后的完整脚本如下:

import boto3
from datetime import datetime
import os

# 初始化S3客户端
session_root = boto3.Session(region_name='eu-west-3', profile_name='my_profile')
s3_client = session_root.client('s3') 

bucket_name = 'my_bucket' 
prefix = 'object1/object2/object3/object4/'

# 1. 遍历所有对象,按目录分组并记录每个目录的最新修改时间
directory_time_map = {}
response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=prefix)

while True:
    contents = response.get('Contents', [])
    for obj in contents:
        # 跳过前缀本身的空对象
        if obj['Key'] == prefix:
            continue
        # 提取对象所属的日期目录
        relative_path = obj['Key'][len(prefix):]
        directory_name = prefix + relative_path.split('/')[0] + '/'
        # 更新目录的最新修改时间
        last_modified = obj['LastModified']
        if directory_name not in directory_time_map or last_modified > directory_time_map[directory_name]:
            directory_time_map[directory_name] = last_modified
    
    # 处理分页(对象数量超过1000时触发)
    if not response.get('IsTruncated'):
        break
    response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=prefix, ContinuationToken=response['NextContinuationToken'])

# 2. 确定最新的目标目录
if not directory_time_map:
    print("未找到任何子目录")
    exit()
latest_directory = max(directory_time_map, key=directory_time_map.get)
print("最新目录:", latest_directory)

# 3. 批量下载该目录下的所有文件
response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=latest_directory)

while True:
    contents = response.get('Contents', [])
    for obj in contents:
        # 跳过S3模拟目录的空对象
        if obj['Size'] == 0:
            continue
        # 构建本地保存路径,保持与S3一致的目录结构
        local_path = obj['Key'].replace(prefix, '')
        local_dir = os.path.dirname(local_path)
        # 创建本地目录(不存在则自动创建)
        if not os.path.exists(local_dir):
            os.makedirs(local_dir, exist_ok=True)
        # 下载文件
        s3_client.download_file(bucket_name, obj['Key'], local_path)
        print(f"已下载: {obj['Key']} -> {local_path}")
    
    # 处理分页
    if not response.get('IsTruncated'):
        break
    response = s3_client.list_objects_v2(Bucket=bucket_name, Prefix=latest_directory, ContinuationToken=response['NextContinuationToken'])

print("所有文件下载完成")

关键修改说明

  • 分页处理:添加分页逻辑,避免因S3返回对象数量超过1000条导致的内容遗漏。
  • 目录分组:通过分割对象Key提取所属目录,确保每个目录的最新时间是其下所有文件的最晚修改时间。
  • 本地目录同步:自动创建与S3对应的本地目录结构,避免文件保存路径错误。
  • 过滤空对象:跳过S3中模拟目录的空对象,只下载实际文件。

内容的提问来源于stack exchange,提问作者mlflow

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 07:11:10