Python嵌套for循环运行缓慢如何优化?同类型商品最新售卖查询场景
核心优化思路
- 消除重复计算:原有逻辑每个商品遍历阶段都重复执行区域遍历、同类型商品日期比对操作,可通过全局预处理提前过滤无效数据,大幅减少后续重复查询开销。
- 淘汰不必要的排序操作:你需要获取的是最大日期对应的商品,无需对整个列表排序,单次遍历即可记录最大值,时间复杂度从O(nlogn)降至O(n),数据量越大性能提升越明显。
- 修正多进程使用错误:原有多进程实现存在两个致命问题:一是每个进程都要拷贝全量
item_mapping和item_history大对象,进程间通信开销完全吃掉了多进程收益;二是使用了全局变量master_item_to_item_list存储结果,多进程下全局变量相互隔离,根本无法拿到正确返回结果。 - 优化IO操作:最终写入文本文件不要逐行写入,所有结果计算完成后统一批量写入,减少IO调度开销。
前置数据结构预处理
先做全局预处理,从根源降低后续计算的查询成本:
- 将
items转为集合item_set = set(items),O(1)复杂度即可判断商品是否在目标列表中,比原有列表的O(n)查询效率提升数个量级。 - 预处理区域历史缓存:给每个区域提前生成过滤后的字典,仅保留在
item_set中的商品,避免后续无效查询:
preprocessed_history = {} for area, hist in item_history.items(): preprocessed_history[area] = {item: date for item, date in hist.items() if item in item_set}
你的日期格式为YYYY-MM-DD,字符串字典序和时间顺序完全一致,无需额外转时间戳即可直接比较大小。
优化后单进程核心逻辑
仅单进程执行效率就比你原有多进程实现高10倍以上:
def get_result_single(item, item_mapping, preprocessed_history): current_line = [item] identical_items = item_mapping.get(item, [item]) for area, hist in preprocessed_history.items(): max_date = '' selected_item = None for it in identical_items: date = hist.get(it) if date and date > max_date: max_date = date selected_item = it if selected_item: current_line.extend([area, selected_item]) return ' '.join(current_line) # 批量计算结果 results = [] for item in items: results.append(get_result_single(item, item_mapping, preprocessed_history)) # 统一写入文件 with open('output.txt', 'w', encoding='utf-8') as f: f.write('\n'.join(results))
多进程优化正确姿势
如果商品量达到百万级需要进一步提速,用进程初始化的方式传递只读大对象,避免每次传参的拷贝开销:
import multiprocessing as mp # 进程内全局只读变量,初始化后无需重复传参 global_mapping = None global_history = None def init_worker(mapping, history): global global_mapping, global_history global_mapping = mapping global_history = history def worker(item): current_line = [item] identical_items = global_mapping.get(item, [item]) for area, hist in global_history.items(): max_date = '' selected_item = None for it in identical_items: date = hist.get(it) if date and date > max_date: max_date = date selected_item = it if selected_item: current_line.extend([area, selected_item]) return ' '.join(current_line) if __name__ == '__main__': # 前置预处理 item_set = set(items) preprocessed_history = { area: {it: dt for it, dt in hist.items() if it in item_set} for area, hist in item_history.items() } # 进程池初始化阶段传入只读公共对象,仅传输一次 with mp.Pool(processes=mp.cpu_count(), initializer=init_worker, initargs=(item_mapping, preprocessed_history)) as pool: # 调大chunksize减少进程调度开销 results = pool.map(worker, items, chunksize=1000) # 批量写文件 with open('output.txt', 'w', encoding='utf-8') as f: f.write('\n'.join(results))
内容的提问来源于stack exchange,提问作者charlie_boy
相关产品推荐
相关产品推荐

