如何使用Boto3对Amazon S3执行查询筛选非指定存储类的对象
AWS CLI 中的 --query 参数是 CLI 层内置的 JMESPath 结果筛选能力,不属于 S3 原生 API 的请求参数,因此 boto3 官方没有提供对应入参,你可以通过以下两种方式实现完全一致的筛选效果:
方案一:手动筛选(无额外依赖)
不需要安装第三方库,直接对 boto3 返回的结果做过滤,同时兼容分页场景和默认存储类的字段兼容问题:
import boto3 s3 = boto3.client('s3') bucket_name = "替换为你的存储桶名称" target_class = "DEEP_ARCHIVE" # 分页器适配对象数量超过1000的场景 paginator = s3.get_paginator('list_objects_v2') result_pages = paginator.paginate(Bucket=bucket_name) unmatched_objs = [] for page in result_pages: if 'Contents' not in page: continue for obj in page['Contents']: # 未返回StorageClass的对象默认属于STANDARD标准存储 storage_class = obj.get('StorageClass', 'STANDARD') if storage_class != target_class: unmatched_objs.append(obj) print(unmatched_objs)
方案二:复用原有JMESPath查询逻辑
如果你希望直接复用 CLI 中的查询语句,可以借助 JMESPath 官方 Python 库实现,语法和 CLI 的 --query 参数完全一致:
- 先安装依赖:
pip install jmespath
- 调用示例:
import boto3 import jmespath s3 = boto3.client('s3') bucket_name = "替换为你的存储桶名称" paginator = s3.get_paginator('list_objects_v2') result_pages = paginator.paginate(Bucket=bucket_name) # 直接使用原有查询语句 query = jmespath.compile('Contents[?StorageClass!=`DEEP_ARCHIVE`]') unmatched_objs = [] for page in result_pages: page_res = query.search(page) if page_res: unmatched_objs.extend(page_res) print(unmatched_objs)
提示:推荐使用
list_objects_v2替代旧版list_objects接口,后者已被AWS官方标记为 Legacy 接口,不建议在新代码中使用。
内容的提问来源于stack exchange,提问作者Frank
相关产品推荐
相关产品推荐

