使用boto3加速AWS S3桶属性数据获取的方法
问题:加速跨环境获取AWS S3桶元数据的脚本
我需要获取AWS账户下所有S3桶的属性数据,现有代码可正常运行(在桶所在环境外执行),但耗时极长。由于找不到包含所需全部元数据的可检索对象,我使用了多种不同的获取方法。目前已减少每个桶的对象检索数量,但每个桶仍耗时超1秒,寻求提速方案。
原代码如下:
import boto3 import csv import time from botocore.exceptions import ClientError s3 = boto3.resource('s3') s3_client = boto3.client('s3') def getTagContents(tagging, keyword): storage = list(filter(lambda tag: tag['Key'] == keyword, tagging)) if storage != "Unknown": if len(storage) > 0: storage = storage[0]['Value'] return storage start_time = time.time() with open("quicksight-test.csv", "w", newline='') as csv_file: fieldnames = ["Name", 'CreationDate', 'Versioned', 'Region', 'Product', 'ProductComponent', 'Environment', 'CustomerName', 'CustomerState', 'StorageClass'] csv_writer = csv.writer(csv_file) csv_writer.writerow(fieldnames) response = s3.buckets.limit(50) for bucket in response: try: tagging = s3_client.get_bucket_tagging(Bucket = bucket.name)['TagSet'] except ClientError as e: continue sclass = list(s3.Bucket(bucket.name).objects.limit(1)) if len(sclass) > 0: sclass = sclass[0].storage_class csv_writer.writerow([bucket.name, bucket.creation_date, bucket.Versioning().status,s3_client.get_bucket_location(Bucket = bucket.name)['LocationConstraint'],getTagContents(tagging, "Project"), getTagContents(tagging, "ProductComponent")," ",getTagContents(tagging, "CustomerName"), getTagContents(tagging, "CustomerState"), sclass]) end_time = time.time() - start_time print(end_time)
提速优化方案
1. 用并行请求替代串行处理
当前脚本逐个桶串行调用API是最大耗时点,使用concurrent.futures.ThreadPoolExecutor实现多线程并行请求,同时根据AWS API速率限制设置合理线程数(建议10-20),可大幅减少总耗时。
2. 减少无效API调用与资源开销
- 替换
bucket.Versioning().status为s3_client.get_bucket_versioning(Bucket=bucket.name),避免每次创建新的Versioning对象,降低资源消耗。 - 用客户端调用
s3_client.list_objects_v2(Bucket=bucket.name, MaxKeys=1)替代资源类调用s3.Bucket(bucket.name).objects.limit(1),客户端调用更轻量高效。
3. 优化标签处理逻辑
将标签列表转换为字典,避免每次调用getTagContents都遍历整个列表:
tag_dict = {tag['Key']: tag['Value'] for tag in tagging_resp['TagSet']}
之后直接用tag_dict.get('TagKey', '')获取对应值,效率远高于filter遍历。
4. 完善错误处理与默认值
- 捕获
get_bucket_tagging异常时,不要直接跳过桶,而是给标签字段设空值,保留无标签桶的其他元数据。 - 处理空桶场景,直接给存储类设默认值(如
STANDARD),避免无效判断。 - 处理
get_bucket_location返回None的情况(如us-east-1区域),统一转换为字符串us-east-1。
优化后的代码示例
import boto3 import csv import time from concurrent.futures import ThreadPoolExecutor, as_completed from botocore.exceptions import ClientError s3_client = boto3.client('s3') def get_bucket_metadata(bucket_name): metadata = { 'Name': bucket_name, 'CreationDate': '', 'Versioned': 'Disabled', 'Region': '', 'Product': '', 'ProductComponent': '', 'Environment': '', 'CustomerName': '', 'CustomerState': '', 'StorageClass': 'STANDARD' } try: # 获取桶创建时间 bucket = boto3.resource('s3').Bucket(bucket_name) metadata['CreationDate'] = bucket.creation_date # 获取版本控制状态 versioning_resp = s3_client.get_bucket_versioning(Bucket=bucket_name) metadata['Versioned'] = versioning_resp.get('Status', 'Disabled') # 获取桶区域 location_resp = s3_client.get_bucket_location(Bucket=bucket_name) metadata['Region'] = location_resp['LocationConstraint'] or 'us-east-1' # 获取标签 try: tagging_resp = s3_client.get_bucket_tagging(Bucket=bucket_name) tag_dict = {tag['Key']: tag['Value'] for tag in tagging_resp['TagSet']} metadata['Product'] = tag_dict.get('Project', '') metadata['ProductComponent'] = tag_dict.get('ProductComponent', '') metadata['CustomerName'] = tag_dict.get('CustomerName', '') metadata['CustomerState'] = tag_dict.get('CustomerState', '') except ClientError: pass # 获取存储类 try: objects_resp = s3_client.list_objects_v2(Bucket=bucket_name, MaxKeys=1) if 'Contents' in objects_resp and objects_resp['Contents']: metadata['StorageClass'] = objects_resp['Contents'][0]['StorageClass'] except ClientError: metadata['StorageClass'] = 'Unknown' except ClientError as e: print(f"处理桶 {bucket_name} 出错: {e}") return metadata start_time = time.time() # 获取所有桶列表 buckets = [bucket.name for bucket in boto3.resource('s3').buckets.all()] # 并行处理并写入CSV with open("quicksight-test.csv", "w", newline='') as csv_file: fieldnames = ["Name", 'CreationDate', 'Versioned', 'Region', 'Product', 'ProductComponent', 'Environment', 'CustomerName', 'CustomerState', 'StorageClass'] csv_writer = csv.DictWriter(csv_file, fieldnames=fieldnames) csv_writer.writeheader() # 线程数可根据实际情况调整 with ThreadPoolExecutor(max_workers=15) as executor: futures = {executor.submit(get_bucket_metadata, bucket): bucket for bucket in buckets} for future in as_completed(futures): csv_writer.writerow(future.result()) end_time = time.time() - start_time print(f"总耗时: {end_time:.2f} 秒")
内容的提问来源于stack exchange,提问作者Gordon Moore
相关产品推荐
相关产品推荐

