如何基于公共id高效合并多个字典列表以适配大数据量处理场景
优化方案
核心优化逻辑
- 原实现的四层嵌套循环时间复杂度为
O(LEVELS长度 * len(B) * len(C) * len(A)),数据量增长后耗时会指数级上升 - 利用id全局唯一的特性,用空间换时间,将A、C列表预转为以
id为键的哈希字典,单次id匹配耗时从O(n)降至O(1) - 对B列表按
level字段预分组,保证输出顺序和LEVELS列表要求一致
实现代码
from collections import defaultdict LEVELS = ["wind", "water", "ice", "earth"] A = [{"id": 4, "active": True}, {"id": 5, "active": True}, {"id": 7, "active": False}] B = [{"id": 5, "status": "red", "level": "water"}, {"id": 4, "status": "green", "level": "wind"}, {"id": 7, "status": "yellow", "level": "wind"}] C = [{"id": 7, "value": 5}, {"id": 4, "value": None}, {"id": 5, "value": 8}] # 预构建id映射字典 a_map = {item["id"]: item for item in A} c_map = {item["id"]: item for item in C} # 预按level分组B的元素 b_level_group = defaultdict(list) for item in B: b_level_group[item["level"]].append(item) # 按LEVELS顺序遍历输出 for level in LEVELS: if level not in b_level_group: continue for b_item in b_level_group[level]: current_id = b_item["id"] print(f"Level: {level} - {current_id} - {a_map[current_id]['active']} - {b_item['status']} - {c_map[current_id]['value']}")
效果说明
优化后整体时间复杂度为线性的O(len(A) + len(B) + len(C)),可稳定处理十万甚至百万级的数据量,输出结果和预期完全一致:
Level: wind - 4 - True - green - None Level: wind - 7 - False - yellow - 5 Level: water - 5 - True - red - 8
内容的提问来源于stack exchange,提问作者xtlc
相关产品推荐
相关产品推荐

