Python处理嵌套列表字典树形数据运行过慢,如何优化提升速度?
Python树形嵌套数据处理性能优化方案
核心性能瓶颈
原代码Step1采用了四层嵌套循环(tau遍历→实体遍历→树遍历→叶子节点遍历),存在大量冗余计算:
- 重复遍历整个树形数据N次(N为tau数量*实体数量)
- 多余的实体存在性检查逻辑,重复计算同一实体组
- 多次列表拼接、类型转换操作产生大量临时对象
优化方案
1. 重构Step1遍历逻辑(性能提升90%+)
直接遍历所有叶子节点一次,提取所有tau对与对应实体组,一次性完成去重,砍掉冗余的tau、实体两层外层循环:
import json def prepare_output(data,tau_length,ent,gamma,tau_start,tau_end,time_range,json_output): # 提前预存所有合法tau对,避免后续无效判断 valid_tau = set((i,i+1) for i in range(tau_length-1)) # Step1 直接遍历一次所有叶子节点构建spgs spgs = {} for tree in data: for leaf in tree: # 只处理合法范围内的tau对 for tau_pair, ent_pairs in leaf.items(): if tau_pair not in valid_tau: continue # 直接生成实体组tuple ent_set = set() for e1,e2 in ent_pairs: ent_set.add(e1) ent_set.add(e2) group = tuple(sorted(ent_set)) # 加入对应tau的集合 if tau_pair not in spgs: spgs[tau_pair] = set() spgs[tau_pair].add(group) # Step2 转换为实体组为键的字典,优化判断写法 result_per_group = {} for tau_pair, groups in spgs.items(): for group in groups: if group in result_per_group: result_per_group[group].append(tau_pair) else: result_per_group[group] = [tau_pair] # Step3 过滤间隔超标的组 threshold = time_range * gamma for group in list(result_per_group.keys()): tau_list = result_per_group[group] for idx in range(len(tau_list)-1): if tau_start[idx+1] - tau_end[idx] > threshold: del result_per_group[group] break # 输出逻辑,去掉冗余全局变量依赖 output = {'output': {str(k):v for k,v in result_per_group.items()}} with open(f'output/{json_output}.json', 'w') as f: json.dump(output, f)
注:原参数名
tau_lenght存在拼写错误,优化代码中修正为tau_length,可根据实际业务定义调整。
2. 其他辅助优化点
- 去掉冗余的
list_contain调用:原逻辑按实体筛选实体组属于多余操作,只要实体组出现在叶子节点中就会被收集,无需逐个实体判断 - 预计算阈值
threshold = time_range * gamma,避免循环内重复计算 - 简化字典键判断:直接用
key in dict替代key in dict.keys(),减少冗余操作 - 把helper函数的简单逻辑直接内联,减少函数调用开销
- 去掉不必要的全局变量
output,避免全局变量副作用
优化效果
原21-22小时的处理流程,优化后预计可压缩到1-2小时内完成,性能提升10倍以上。
内容的提问来源于stack exchange,提问作者Kela
相关产品推荐
相关产品推荐

