You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS新手求助:如何检查S3文件是否存在并仅读取存在的文件?

解决S3文件不存在时跳过读取的问题

针对你需要每日读取指定格式S3文件、不存在则跳过的需求,下面提供几种实用的实现方案:

方案一:先检查文件存在性再读取

通过boto3的head_object方法先验证文件是否存在,存在再执行读取操作:

import boto3
from botocore.exceptions import ClientError

# 初始化S3客户端
s3 = boto3.client('s3')
# 替换为你的存储桶名称
bucket_name = 'your-bucket-name'

def check_s3_file_exists(bucket, file_key):
    try:
        # 发送HEAD请求验证文件存在
        s3.head_object(Bucket=bucket, Key=file_key)
        return True
    except ClientError as e:
        # 404状态码表示文件不存在
        if e.response['Error']['Code'] == '404':
            return False
        # 其他异常(如权限错误)正常抛出
        raise

# 替换为你需要处理的日期列表,可根据实际逻辑自动生成
target_dates = ['2022-10-01', '2022-10-02', '2022-10-03']

for date in target_dates:
    file_key = f'abc{date}'
    if check_s3_file_exists(bucket_name, file_key):
        read_from_s3(file_key)
    else:
        print(f"文件 {file_key} 不存在,跳过")

方案二:捕获读取时的异常(推荐)

遵循Python“请求原谅比许可容易”的设计哲学,直接尝试读取文件,捕获404异常来处理文件不存在的情况,能避免检查与读取之间的竞态问题(比如检查后文件被删除):

import boto3
from botocore.exceptions import ClientError

s3 = boto3.client('s3')
bucket_name = 'your-bucket-name'
target_dates = ['2022-10-01', '2022-10-02', '2022-10-03']

def read_from_s3(file_key):
    try:
        # 执行S3文件读取逻辑
        response = s3.get_object(Bucket=bucket_name, Key=file_key)
        file_content = response['Body'].read().decode('utf-8')
        # 这里添加你的后续操作
        print(f"成功读取文件 {file_key}")
    except ClientError as e:
        if e.response['Error']['Code'] == '404':
            print(f"文件 {file_key} 不存在,跳过")
        else:
            # 非404异常(如权限不足)抛出,便于排查问题
            raise

for date in target_dates:
    file_key = f'abc{date}'
    read_from_s3(file_key)

方案三:批量列出文件后匹配(适合大量文件场景)

如果需要处理的文件数量较多,先批量列出存储桶内的所有文件,再与目标文件名匹配,减少API调用次数,提升效率:

import boto3

s3 = boto3.client('s3')
bucket_name = 'your-bucket-name'
target_dates = ['2022-10-01', '2022-10-02', '2022-10-03']

# 批量获取存储桶内所有文件的Key
response = s3.list_objects_v2(Bucket=bucket_name)
existing_file_keys = {obj['Key'] for obj in response.get('Contents', [])}

# 处理分页(如果文件超过1000个)
while response.get('IsTruncated'):
    response = s3.list_objects_v2(Bucket=bucket_name, ContinuationToken=response['NextContinuationToken'])
    existing_file_keys.update({obj['Key'] for obj in response.get('Contents', [])})

for date in target_dates:
    file_key = f'abc{date}'
    if file_key in existing_file_keys:
        read_from_s3(file_key)
    else:
        print(f"文件 {file_key} 不存在,跳过")

注意事项

  • 确保你的AWS身份拥有s3:GetObject和s3:HeadObject(方案一需要)的权限;
  • 若使用自动生成日期的逻辑,可结合datetime模块动态生成每日目标文件名,无需手动维护日期列表。

内容的提问来源于stack exchange,提问作者vidathri

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 01:01:06