You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python嵌套for循环运行缓慢如何优化?同类型商品最新售卖查询场景

核心优化思路
  • 消除重复计算:原有逻辑每个商品遍历阶段都重复执行区域遍历、同类型商品日期比对操作,可通过全局预处理提前过滤无效数据,大幅减少后续重复查询开销。
  • 淘汰不必要的排序操作:你需要获取的是最大日期对应的商品,无需对整个列表排序,单次遍历即可记录最大值,时间复杂度从O(nlogn)降至O(n),数据量越大性能提升越明显。
  • 修正多进程使用错误:原有多进程实现存在两个致命问题:一是每个进程都要拷贝全量item_mapping和item_history大对象,进程间通信开销完全吃掉了多进程收益;二是使用了全局变量master_item_to_item_list存储结果,多进程下全局变量相互隔离,根本无法拿到正确返回结果。
  • 优化IO操作:最终写入文本文件不要逐行写入,所有结果计算完成后统一批量写入,减少IO调度开销。
前置数据结构预处理

先做全局预处理,从根源降低后续计算的查询成本:

  1. 将items转为集合item_set = set(items),O(1)复杂度即可判断商品是否在目标列表中,比原有列表的O(n)查询效率提升数个量级。
  2. 预处理区域历史缓存:给每个区域提前生成过滤后的字典,仅保留在item_set中的商品,避免后续无效查询:
preprocessed_history = {}
for area, hist in item_history.items():
    preprocessed_history[area] = {item: date for item, date in hist.items() if item in item_set}

你的日期格式为YYYY-MM-DD,字符串字典序和时间顺序完全一致,无需额外转时间戳即可直接比较大小。

优化后单进程核心逻辑

仅单进程执行效率就比你原有多进程实现高10倍以上:

def get_result_single(item, item_mapping, preprocessed_history):
    current_line = [item]
    identical_items = item_mapping.get(item, [item])
    for area, hist in preprocessed_history.items():
        max_date = ''
        selected_item = None
        for it in identical_items:
            date = hist.get(it)
            if date and date > max_date:
                max_date = date
                selected_item = it
        if selected_item:
            current_line.extend([area, selected_item])
    return ' '.join(current_line)

# 批量计算结果
results = []
for item in items:
    results.append(get_result_single(item, item_mapping, preprocessed_history))

# 统一写入文件
with open('output.txt', 'w', encoding='utf-8') as f:
    f.write('\n'.join(results))
多进程优化正确姿势

如果商品量达到百万级需要进一步提速,用进程初始化的方式传递只读大对象,避免每次传参的拷贝开销:

import multiprocessing as mp

# 进程内全局只读变量,初始化后无需重复传参
global_mapping = None
global_history = None

def init_worker(mapping, history):
    global global_mapping, global_history
    global_mapping = mapping
    global_history = history

def worker(item):
    current_line = [item]
    identical_items = global_mapping.get(item, [item])
    for area, hist in global_history.items():
        max_date = ''
        selected_item = None
        for it in identical_items:
            date = hist.get(it)
            if date and date > max_date:
                max_date = date
                selected_item = it
        if selected_item:
            current_line.extend([area, selected_item])
    return ' '.join(current_line)

if __name__ == '__main__':
    # 前置预处理
    item_set = set(items)
    preprocessed_history = {
        area: {it: dt for it, dt in hist.items() if it in item_set}
        for area, hist in item_history.items()
    }
    # 进程池初始化阶段传入只读公共对象,仅传输一次
    with mp.Pool(processes=mp.cpu_count(), initializer=init_worker, initargs=(item_mapping, preprocessed_history)) as pool:
        # 调大chunksize减少进程调度开销
        results = pool.map(worker, items, chunksize=1000)
    # 批量写文件
    with open('output.txt', 'w', encoding='utf-8') as f:
        f.write('\n'.join(results))

内容的提问来源于stack exchange,提问作者charlie_boy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 13:36:05