You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python多线程优化失效:大列表数据匹配效率提升求助

大列表模糊匹配的高效优化方案

问题背景

  • 处理两个大列表:
    • json_list:包含23000+条带code字段的字典数据
    • marking:包含3000+条字符串
  • 需求:基于code字段,用SequenceMatcher匹配(匹配度为1 或 0.98-0.99)筛选数据
  • 现状:单线程耗时超20分钟,现有多线程实现效率仍未达标

现有代码的核心问题

import math
import threading
from difflib import SequenceMatcher

def find_marking(x=None, y=None):
    text_match = SequenceMatcher(None, x, y.get('code')).ratio()
    if text_match == 1 or (0.98 <= text_match < 0.99): return y
    return None

def eliminate_marking(marking_list=None, json_list=None)->tuple:
    result,result_mark = [], []

    def __process_eliminate(marking=None, data_scrap=None):
        for a in range(0, math.ceil(len(data_scrap)/100)):
            if len(data_scrap[a * 100:(a + 1) * 100]) > 0:
                for data in data_scrap[a * 100:(a + 1) * 100]:
                    result_data = find_marking(marking, data)
                    if result_data:
                        data_scrap.pop(0)
                        result_mark.append(marking)
                        result.append(result_data)
                        return

    threads = []
    for marking in marking_list:
        th = threading.Thread(target=__process_eliminate, kwargs={'marking':marking,
                                                                'data_scrap': json_list})
        th.start()
        threads.append(th)
    
    for thread in threads: thread.join()

    return result_mark, result

现有代码存在以下低效/风险点:

  • 线程安全隐患:多线程同时修改json_list、result、result_mark,会导致数据竞争,出现重复匹配、数据丢失或异常
  • 遍历逻辑冗余:每个线程切片遍历整个列表,找到一个匹配就返回,切片操作和循环嵌套增加不必要的开销
  • 线程资源浪费:为3000+条marking各创建一个线程,线程数量远超CPU核心数,上下文切换开销抵消了并行收益
  • 无匹配缓存:同一个code可能被多个marking重复计算匹配度,浪费算力

优化方案

1. 预处理:构建索引缩小候选范围

先对json_list做预处理,按code长度分组,同时构建code到数据的映射,后续只在长度相近的候选集中匹配(匹配度0.98以上要求长度差异极小):

from collections import defaultdict

def preprocess_json_list(json_list):
    code_groups = defaultdict(list)  # key: code长度,value: 对应长度的code列表
    code_to_data = {}                # key: code,value: 完整字典数据
    for data in json_list:
        code = data.get('code')
        if code:
            code_groups[len(code)].append(code)
            code_to_data[code] = data
    return code_groups, code_to_data

def get_candidate_codes(marking, code_groups):
    target_len = len(marking)
    # 仅考虑长度差为0或1的code组,过滤掉不可能达到高匹配度的候选
    candidate_lens = [target_len, target_len + 1, target_len - 1]
    candidates = []
    for length in candidate_lens:
        candidates.extend(code_groups.get(length, []))
    return candidates

2. 优化匹配算法+缓存结果

替换低效匹配逻辑,先做快速过滤,再用精确匹配,同时缓存已计算的匹配结果避免重复计算:

from difflib import SequenceMatcher
from functools import lru_cache

# 缓存(marking, code)的匹配结果,避免重复计算
@lru_cache(maxsize=None)
def is_match(marking, code):
    # 快速长度校验:长度差超过1直接排除
    len_diff = abs(len(marking) - len(code))
    if len_diff > 1:
        return False
    # 精确计算匹配度
    ratio = SequenceMatcher(None, marking, code).ratio()
    return ratio == 1 or (0.98 <= ratio < 0.99)

3. 改用多进程充分利用多核CPU

Python的GIL限制了多线程在CPU密集型任务中的效率,改用multiprocessing进程池,充分发挥多核算力:

import multiprocessing

def process_single_marking(marking, code_groups, code_to_data):
    candidates = get_candidate_codes(marking, code_groups)
    for code in candidates:
        if is_match(marking, code):
            return (marking, code_to_data[code])
    return None

def optimized_eliminate(marking_list, json_list):
    code_groups, code_to_data = preprocess_json_list(json_list)
    
    # 进程池数量设为CPU核心数,避免过度调度
    with multiprocessing.Pool(processes=multiprocessing.cpu_count()) as pool:
        # 批量提交任务
        task_args = [(m, code_groups, code_to_data) for m in marking_list]
        raw_results = pool.starmap(process_single_marking, task_args)
    
    # 过滤无效结果并去重(避免多个marking匹配同一个code)
    seen_codes = set()
    final_marks = []
    final_data = []
    for res in raw_results:
        if res:
            mark, data = res
            code = data.get('code')
            if code not in seen_codes:
                seen_codes.add(code)
                final_marks.append(mark)
                final_data.append(data)
    
    return final_marks, final_data

4. 可选:替换为更高效的模糊匹配库

标准库difflib的SequenceMatcher性能一般,可改用rapidfuzz(比difflib快10-100倍):

from rapidfuzz import fuzz
from functools import lru_cache

@lru_cache(maxsize=None)
def is_match(marking, code):
    len_diff = abs(len(marking) - len(code))
    if len_diff > 1:
        return False
    # rapidfuzz返回0-100的匹配分数,转换为0-1的比例
    ratio = fuzz.ratio(marking, code) / 100
    return ratio == 1.0 or (0.98 <= ratio < 0.99)

内容的提问来源于stack exchange,提问作者dimas setiawan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 18:10:58