You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在AWS Cost Explorer API的get_cost_and_usage中分组超2个字段?

解决AWS Cost Explorer API多分组(双自定义标签+服务)的问题

AWS Cost Explorer的get_cost_and_usage接口确实限制最多2个GroupBy字段,要实现双自定义标签+服务的三维分组,最靠谱的方式就是拆分多轮查询,再用Pandas合并数据。以下是两种可行的实现思路:

一、精准遍历标签组合查询(推荐)

先获取所有自定义标签的组合,再针对每个组合单独查询对应服务的成本,避免合并时的数据歧义:

1. 初始化客户端与通用函数

import boto3
import pandas as pd

ce_client = boto3.client('ce')

def get_tag_combinations(time_period, tag1_key, tag2_key):
    """获取两个自定义标签的所有唯一组合"""
    response = ce_client.get_cost_and_usage(
        TimePeriod=time_period,
        Granularity='MONTHLY',
        Metrics=["UnblendedCost"],
        GroupBy=[
            {'Type': 'TAG', 'Key': tag1_key},
            {'Type': 'TAG', 'Key': tag2_key}
        ]
    )
    combinations = []
    for group in response['ResultsByTime'][0]['Groups']:
        combinations.append({
            tag1_key: group['Keys'][0],
            tag2_key: group['Keys'][1]
        })
    # 去重返回
    return pd.DataFrame(combinations).drop_duplicates()

2. 遍历组合查询服务成本

假设你需要按Project(自定义标签1)、Environment(自定义标签2)、Service分组:

# 按需设置时间范围
time_period = {'Start': '2024-01-01', 'End': '2024-02-01'}
tag1, tag2 = 'Project', 'Environment'

# 获取所有标签组合
tag_df = get_tag_combinations(time_period, tag1, tag2)

# 遍历每个组合查询对应服务成本
final_data = []
for _, row in tag_df.iterrows():
    tag1_val = row[tag1]
    tag2_val = row[tag2]
    
    # 过滤当前标签组合,按服务分组查询
    cost_response = ce_client.get_cost_and_usage(
        TimePeriod=time_period,
        Granularity='MONTHLY',
        Metrics=["UnblendedCost"],
        GroupBy=[{'Type': 'DIMENSION', 'Key': 'SERVICE'}],
        Filter={
            'And': [
                {'Tags': {'Key': tag1, 'Values': [tag1_val]}},
                {'Tags': {'Key': tag2, 'Values': [tag2_val]}}
            ]
        }
    )
    
    # 解析结果并补充标签字段
    for group in cost_response['ResultsByTime'][0]['Groups']:
        final_data.append({
            tag1: tag1_val,
            tag2: tag2_val,
            'Service': group['Keys'][0],
            'Cost': float(group['Metrics']['UnblendedCost']['Amount'])
        })

# 生成最终DataFrame
final_df = pd.DataFrame(final_data)

二、三次查询合并(适合标签组合较多场景)

如果标签组合数量过大,遍历查询效率低,可以分三次查询不同维度组合,再通过Pandas关联合并:

1. 三次基础查询

# 查询1:标签1+服务
df_tag1_service = get_cost_data(
    group_by=[{'Type': 'TAG', 'Key': tag1}, {'Type': 'DIMENSION', 'Key': 'SERVICE'}],
    time_period=time_period
)

# 查询2:标签1+标签2
df_tags = get_cost_data(
    group_by=[{'Type': 'TAG', 'Key': tag1}, {'Type': 'TAG', 'Key': tag2}],
    time_period=time_period
)

# 查询3:标签2+服务
df_tag2_service = get_cost_data(
    group_by=[{'Type': 'TAG', 'Key': tag2}, {'Type': 'DIMENSION', 'Key': 'SERVICE'}],
    time_period=time_period
)

2. 合并与去重

# 先关联标签1+标签2与标签1+服务
merged_df = pd.merge(df_tags, df_tag1_service, on=tag1, how='inner')
# 再关联标签2+服务,得到三维组合
final_df = pd.merge(merged_df, df_tag2_service, on=[tag2, 'SERVICE'], how='inner')

# 按三维分组聚合成本(避免重复数据)
final_df = final_df.groupby([tag1, tag2, 'SERVICE'])['Cost'].sum().reset_index()

三、注意事项

  • 若存在空标签值,合并时用how='outer'避免丢失数据
  • 高频查询可能触发API限流,可添加tenacity库实现重试逻辑
  • 复杂场景推荐直接使用AWS Cost and Usage Report(CUR):导出到S3后用Athena或Pandas直接分析,无GroupBy数量限制

内容的提问来源于stack exchange,提问作者egg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 15:15:31