Python按字典name键求和rows_processed、pipelines并生成指定格式数组
问题解决代码实现
原有代码问题说明
- 未使用
name作为分组统计的key,错误遍历了字典所有键,逻辑偏离需求 - 字典
get方法传入参数为字段值而非分组名,累加逻辑完全错误 - None值处理逻辑未考虑历史累加值,直接赋值0会覆盖之前的统计结果
正确实现代码
版本1(无需导入第三方模块)
rows_sum = {} pipes_sum = {} for item in batches: name = item["name"] # None值统一按0处理 row_add = item["rows_processed"] if item["rows_processed"] is not None else 0 pipe_add = item["pipelines"] if item["pipelines"] is not None else 0 # 按name分组累加 rows_sum[name] = rows_sum.get(name, 0) + row_add pipes_sum[name] = pipes_sum.get(name, 0) + pipe_add # 按name后缀数字降序排列,和示例输出顺序一致 sorted_names = sorted(rows_sum.keys(), key=lambda x: -int(x.split()[-1])) # 组装最终结果 result = { "Rows Processed": [rows_sum[name] for name in sorted_names], "Pipelines Processed": [pipes_sum[name] for name in sorted_names] }
版本2(使用defaultdict简化代码)
from collections import defaultdict rows_sum = defaultdict(int) pipes_sum = defaultdict(int) for item in batches: name = item["name"] rows_sum[name] += item["rows_processed"] if item["rows_processed"] is not None else 0 pipes_sum[name] += item["pipelines"] if item["pipelines"] is not None else 0 sorted_names = sorted(rows_sum.keys(), key=lambda x: -int(x.split()[-1])) result = { "Rows Processed": [rows_sum[name] for name in sorted_names], "Pipelines Processed": [pipes_sum[name] for name in sorted_names] }
内容的提问来源于stack exchange,提问作者Hassan Shahbaz
相关产品推荐
相关产品推荐

