如何用Python itertools groupby实现按环境统计错误类型及数量?
统计每个环境下的错误类型出现次数
示例错误数据
errors = [ {'PartitionKey': '34', 'RowKey': '14', 'Component': 'mamba', 'Environment': 'QA', 'Error': '404 not found', 'Group': 'Test', 'Job': 'cutting', 'JobType': 'automated'}, {'PartitionKey': '35', 'RowKey': '15', 'Component': 'mamba', 'Environment': 'QA', 'Error': '404 not found', 'Group': 'Test', 'Job': 'cutting', 'JobType': 'automated'}, {'PartitionKey': '36', 'RowKey': '16', 'Component': 'mamba', 'Environment': 'Dev', 'Error': '404 not found', 'Group': 'random', 'Job': 'moping', 'JobType': 'manual'}, {'PartitionKey': '37', 'RowKey': '17', 'Component': 'mamba', 'Environment': 'QA', 'Error': '404 not found', 'Group': 'Test', 'Job': 'cutting', 'JobType': 'automated'}, {'PartitionKey': '38', 'RowKey': '18', 'Component': 'mamba', 'Environment': 'Dev', 'Error': '404 not found', 'Group': 'random', 'Job': 'moping', 'JobType': 'manual'}, {'PartitionKey': '39', 'RowKey': '19', 'Component': 'Scorpio', 'Environment': 'Dev', 'Error': '500 internal error', 'Group': 'minerva', 'Job': 'cleaning', 'JobType': 'manual'}, {'PartitionKey': '39', 'RowKey': '19', 'Component': 'Scorpio', 'Environment': 'Dev', 'Error': '500 internal error', 'Group': 'minerva', 'Job': 'cleaning', 'JobType': 'manual'} ]
期望输出格式
{ 'QA': { '404 not found': 3 }, 'Dev': { '404 not found': 2, '500 internal error': 2 } }
你的代码问题分析
你提供的代码存在以下问题:
- 变量名错误:
valss应为vals,group应为val - 逻辑错误:仅将错误对象添加到列表,未实现计数功能
- 结构错误:未构建以环境为外层键、错误类型为内层键的嵌套字典
正确实现方法
方法一:使用defaultdict+Counter(推荐,简洁高效)
from collections import defaultdict, Counter result = defaultdict(Counter) for item in errors: env = item['Environment'] error = item['Error'] result[env][error] += 1 # 转换为普通字典(可选) result_dict = {k: dict(v) for k, v in result.items()} print(result_dict)
方法二:使用itertools.groupby(需先排序)
groupby仅对连续相同的键分组,因此需先按环境和错误类型排序:
from itertools import groupby from collections import defaultdict # 先按环境排序,再按错误类型排序 sorted_errors = sorted(errors, key=lambda x: (x['Environment'], x['Error'])) result = defaultdict(dict) for env, env_group in groupby(sorted_errors, key=lambda x: x['Environment']): for error, error_group in groupby(env_group, key=lambda x: x['Error']): result[env][error] = len(list(error_group)) print(dict(result))
内容的提问来源于stack exchange,提问作者Hound
相关产品推荐
相关产品推荐

