You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用boto3加速AWS S3桶属性数据获取的方法

问题:加速跨环境获取AWS S3桶元数据的脚本

我需要获取AWS账户下所有S3桶的属性数据,现有代码可正常运行(在桶所在环境外执行),但耗时极长。由于找不到包含所需全部元数据的可检索对象,我使用了多种不同的获取方法。目前已减少每个桶的对象检索数量,但每个桶仍耗时超1秒,寻求提速方案。

原代码如下:

import boto3
import csv
import time
from botocore.exceptions import ClientError

s3 = boto3.resource('s3')
s3_client = boto3.client('s3')

def getTagContents(tagging, keyword):
    storage = list(filter(lambda tag: tag['Key'] == keyword, tagging))
    if storage != "Unknown":
        if len(storage) > 0:
            storage = storage[0]['Value']
    return storage

start_time = time.time()
with open("quicksight-test.csv", "w", newline='') as csv_file:
    fieldnames = ["Name", 'CreationDate', 'Versioned', 'Region', 'Product', 'ProductComponent', 'Environment', 'CustomerName', 'CustomerState', 'StorageClass']
    
    csv_writer = csv.writer(csv_file)
    csv_writer.writerow(fieldnames)
    
    response = s3.buckets.limit(50)
    for bucket in response:
        try:
            tagging = s3_client.get_bucket_tagging(Bucket = bucket.name)['TagSet']
        except ClientError as e:
            continue
        
        sclass = list(s3.Bucket(bucket.name).objects.limit(1))
        if len(sclass) > 0:
            sclass = sclass[0].storage_class 
        csv_writer.writerow([bucket.name, bucket.creation_date, bucket.Versioning().status,s3_client.get_bucket_location(Bucket = bucket.name)['LocationConstraint'],getTagContents(tagging, "Project"), getTagContents(tagging, "ProductComponent")," ",getTagContents(tagging, "CustomerName"), getTagContents(tagging, "CustomerState"), sclass])
        
end_time = time.time() - start_time
print(end_time)

提速优化方案

1. 用并行请求替代串行处理

当前脚本逐个桶串行调用API是最大耗时点,使用concurrent.futures.ThreadPoolExecutor实现多线程并行请求,同时根据AWS API速率限制设置合理线程数(建议10-20),可大幅减少总耗时。

2. 减少无效API调用与资源开销

  • 替换bucket.Versioning().status为s3_client.get_bucket_versioning(Bucket=bucket.name),避免每次创建新的Versioning对象,降低资源消耗。
  • 用客户端调用s3_client.list_objects_v2(Bucket=bucket.name, MaxKeys=1)替代资源类调用s3.Bucket(bucket.name).objects.limit(1),客户端调用更轻量高效。

3. 优化标签处理逻辑

将标签列表转换为字典,避免每次调用getTagContents都遍历整个列表:

tag_dict = {tag['Key']: tag['Value'] for tag in tagging_resp['TagSet']}

之后直接用tag_dict.get('TagKey', '')获取对应值,效率远高于filter遍历。

4. 完善错误处理与默认值

  • 捕获get_bucket_tagging异常时,不要直接跳过桶,而是给标签字段设空值,保留无标签桶的其他元数据。
  • 处理空桶场景,直接给存储类设默认值(如STANDARD),避免无效判断。
  • 处理get_bucket_location返回None的情况(如us-east-1区域),统一转换为字符串us-east-1。

优化后的代码示例

import boto3
import csv
import time
from concurrent.futures import ThreadPoolExecutor, as_completed
from botocore.exceptions import ClientError

s3_client = boto3.client('s3')

def get_bucket_metadata(bucket_name):
    metadata = {
        'Name': bucket_name,
        'CreationDate': '',
        'Versioned': 'Disabled',
        'Region': '',
        'Product': '',
        'ProductComponent': '',
        'Environment': '',
        'CustomerName': '',
        'CustomerState': '',
        'StorageClass': 'STANDARD'
    }
    
    try:
        # 获取桶创建时间
        bucket = boto3.resource('s3').Bucket(bucket_name)
        metadata['CreationDate'] = bucket.creation_date
        
        # 获取版本控制状态
        versioning_resp = s3_client.get_bucket_versioning(Bucket=bucket_name)
        metadata['Versioned'] = versioning_resp.get('Status', 'Disabled')
        
        # 获取桶区域
        location_resp = s3_client.get_bucket_location(Bucket=bucket_name)
        metadata['Region'] = location_resp['LocationConstraint'] or 'us-east-1'
        
        # 获取标签
        try:
            tagging_resp = s3_client.get_bucket_tagging(Bucket=bucket_name)
            tag_dict = {tag['Key']: tag['Value'] for tag in tagging_resp['TagSet']}
            metadata['Product'] = tag_dict.get('Project', '')
            metadata['ProductComponent'] = tag_dict.get('ProductComponent', '')
            metadata['CustomerName'] = tag_dict.get('CustomerName', '')
            metadata['CustomerState'] = tag_dict.get('CustomerState', '')
        except ClientError:
            pass
        
        # 获取存储类
        try:
            objects_resp = s3_client.list_objects_v2(Bucket=bucket_name, MaxKeys=1)
            if 'Contents' in objects_resp and objects_resp['Contents']:
                metadata['StorageClass'] = objects_resp['Contents'][0]['StorageClass']
        except ClientError:
            metadata['StorageClass'] = 'Unknown'
            
    except ClientError as e:
        print(f"处理桶 {bucket_name} 出错: {e}")
    
    return metadata

start_time = time.time()

# 获取所有桶列表
buckets = [bucket.name for bucket in boto3.resource('s3').buckets.all()]

# 并行处理并写入CSV
with open("quicksight-test.csv", "w", newline='') as csv_file:
    fieldnames = ["Name", 'CreationDate', 'Versioned', 'Region', 'Product', 'ProductComponent', 'Environment', 'CustomerName', 'CustomerState', 'StorageClass']
    csv_writer = csv.DictWriter(csv_file, fieldnames=fieldnames)
    csv_writer.writeheader()
    
    # 线程数可根据实际情况调整
    with ThreadPoolExecutor(max_workers=15) as executor:
        futures = {executor.submit(get_bucket_metadata, bucket): bucket for bucket in buckets}
        
        for future in as_completed(futures):
            csv_writer.writerow(future.result())

end_time = time.time() - start_time
print(f"总耗时: {end_time:.2f} 秒")

内容的提问来源于stack exchange,提问作者Gordon Moore

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 04:00:06