AWS新手求助:如何检查S3文件是否存在并仅读取存在的文件?
解决S3文件不存在时跳过读取的问题
针对你需要每日读取指定格式S3文件、不存在则跳过的需求,下面提供几种实用的实现方案:
方案一:先检查文件存在性再读取
通过boto3的head_object方法先验证文件是否存在,存在再执行读取操作:
import boto3 from botocore.exceptions import ClientError # 初始化S3客户端 s3 = boto3.client('s3') # 替换为你的存储桶名称 bucket_name = 'your-bucket-name' def check_s3_file_exists(bucket, file_key): try: # 发送HEAD请求验证文件存在 s3.head_object(Bucket=bucket, Key=file_key) return True except ClientError as e: # 404状态码表示文件不存在 if e.response['Error']['Code'] == '404': return False # 其他异常(如权限错误)正常抛出 raise # 替换为你需要处理的日期列表,可根据实际逻辑自动生成 target_dates = ['2022-10-01', '2022-10-02', '2022-10-03'] for date in target_dates: file_key = f'abc{date}' if check_s3_file_exists(bucket_name, file_key): read_from_s3(file_key) else: print(f"文件 {file_key} 不存在,跳过")
方案二:捕获读取时的异常(推荐)
遵循Python“请求原谅比许可容易”的设计哲学,直接尝试读取文件,捕获404异常来处理文件不存在的情况,能避免检查与读取之间的竞态问题(比如检查后文件被删除):
import boto3 from botocore.exceptions import ClientError s3 = boto3.client('s3') bucket_name = 'your-bucket-name' target_dates = ['2022-10-01', '2022-10-02', '2022-10-03'] def read_from_s3(file_key): try: # 执行S3文件读取逻辑 response = s3.get_object(Bucket=bucket_name, Key=file_key) file_content = response['Body'].read().decode('utf-8') # 这里添加你的后续操作 print(f"成功读取文件 {file_key}") except ClientError as e: if e.response['Error']['Code'] == '404': print(f"文件 {file_key} 不存在,跳过") else: # 非404异常(如权限不足)抛出,便于排查问题 raise for date in target_dates: file_key = f'abc{date}' read_from_s3(file_key)
方案三:批量列出文件后匹配(适合大量文件场景)
如果需要处理的文件数量较多,先批量列出存储桶内的所有文件,再与目标文件名匹配,减少API调用次数,提升效率:
import boto3 s3 = boto3.client('s3') bucket_name = 'your-bucket-name' target_dates = ['2022-10-01', '2022-10-02', '2022-10-03'] # 批量获取存储桶内所有文件的Key response = s3.list_objects_v2(Bucket=bucket_name) existing_file_keys = {obj['Key'] for obj in response.get('Contents', [])} # 处理分页(如果文件超过1000个) while response.get('IsTruncated'): response = s3.list_objects_v2(Bucket=bucket_name, ContinuationToken=response['NextContinuationToken']) existing_file_keys.update({obj['Key'] for obj in response.get('Contents', [])}) for date in target_dates: file_key = f'abc{date}' if file_key in existing_file_keys: read_from_s3(file_key) else: print(f"文件 {file_key} 不存在,跳过")
注意事项
- 确保你的AWS身份拥有
s3:GetObject和s3:HeadObject(方案一需要)的权限; - 若使用自动生成日期的逻辑,可结合
datetime模块动态生成每日目标文件名,无需手动维护日期列表。
内容的提问来源于stack exchange,提问作者vidathri
相关产品推荐
相关产品推荐

