Python多线程优化失效:大列表数据匹配效率提升求助
大列表模糊匹配的高效优化方案
问题背景
- 处理两个大列表:
json_list:包含23000+条带code字段的字典数据marking:包含3000+条字符串
- 需求:基于
code字段,用SequenceMatcher匹配(匹配度为1 或 0.98-0.99)筛选数据 - 现状:单线程耗时超20分钟,现有多线程实现效率仍未达标
现有代码的核心问题
import math import threading from difflib import SequenceMatcher def find_marking(x=None, y=None): text_match = SequenceMatcher(None, x, y.get('code')).ratio() if text_match == 1 or (0.98 <= text_match < 0.99): return y return None def eliminate_marking(marking_list=None, json_list=None)->tuple: result,result_mark = [], [] def __process_eliminate(marking=None, data_scrap=None): for a in range(0, math.ceil(len(data_scrap)/100)): if len(data_scrap[a * 100:(a + 1) * 100]) > 0: for data in data_scrap[a * 100:(a + 1) * 100]: result_data = find_marking(marking, data) if result_data: data_scrap.pop(0) result_mark.append(marking) result.append(result_data) return threads = [] for marking in marking_list: th = threading.Thread(target=__process_eliminate, kwargs={'marking':marking, 'data_scrap': json_list}) th.start() threads.append(th) for thread in threads: thread.join() return result_mark, result
现有代码存在以下低效/风险点:
- 线程安全隐患:多线程同时修改
json_list、result、result_mark,会导致数据竞争,出现重复匹配、数据丢失或异常 - 遍历逻辑冗余:每个线程切片遍历整个列表,找到一个匹配就返回,切片操作和循环嵌套增加不必要的开销
- 线程资源浪费:为3000+条marking各创建一个线程,线程数量远超CPU核心数,上下文切换开销抵消了并行收益
- 无匹配缓存:同一个
code可能被多个marking重复计算匹配度,浪费算力
优化方案
1. 预处理:构建索引缩小候选范围
先对json_list做预处理,按code长度分组,同时构建code到数据的映射,后续只在长度相近的候选集中匹配(匹配度0.98以上要求长度差异极小):
from collections import defaultdict def preprocess_json_list(json_list): code_groups = defaultdict(list) # key: code长度,value: 对应长度的code列表 code_to_data = {} # key: code,value: 完整字典数据 for data in json_list: code = data.get('code') if code: code_groups[len(code)].append(code) code_to_data[code] = data return code_groups, code_to_data def get_candidate_codes(marking, code_groups): target_len = len(marking) # 仅考虑长度差为0或1的code组,过滤掉不可能达到高匹配度的候选 candidate_lens = [target_len, target_len + 1, target_len - 1] candidates = [] for length in candidate_lens: candidates.extend(code_groups.get(length, [])) return candidates
2. 优化匹配算法+缓存结果
替换低效匹配逻辑,先做快速过滤,再用精确匹配,同时缓存已计算的匹配结果避免重复计算:
from difflib import SequenceMatcher from functools import lru_cache # 缓存(marking, code)的匹配结果,避免重复计算 @lru_cache(maxsize=None) def is_match(marking, code): # 快速长度校验:长度差超过1直接排除 len_diff = abs(len(marking) - len(code)) if len_diff > 1: return False # 精确计算匹配度 ratio = SequenceMatcher(None, marking, code).ratio() return ratio == 1 or (0.98 <= ratio < 0.99)
3. 改用多进程充分利用多核CPU
Python的GIL限制了多线程在CPU密集型任务中的效率,改用multiprocessing进程池,充分发挥多核算力:
import multiprocessing def process_single_marking(marking, code_groups, code_to_data): candidates = get_candidate_codes(marking, code_groups) for code in candidates: if is_match(marking, code): return (marking, code_to_data[code]) return None def optimized_eliminate(marking_list, json_list): code_groups, code_to_data = preprocess_json_list(json_list) # 进程池数量设为CPU核心数,避免过度调度 with multiprocessing.Pool(processes=multiprocessing.cpu_count()) as pool: # 批量提交任务 task_args = [(m, code_groups, code_to_data) for m in marking_list] raw_results = pool.starmap(process_single_marking, task_args) # 过滤无效结果并去重(避免多个marking匹配同一个code) seen_codes = set() final_marks = [] final_data = [] for res in raw_results: if res: mark, data = res code = data.get('code') if code not in seen_codes: seen_codes.add(code) final_marks.append(mark) final_data.append(data) return final_marks, final_data
4. 可选:替换为更高效的模糊匹配库
标准库difflib的SequenceMatcher性能一般,可改用rapidfuzz(比difflib快10-100倍):
from rapidfuzz import fuzz from functools import lru_cache @lru_cache(maxsize=None) def is_match(marking, code): len_diff = abs(len(marking) - len(code)) if len_diff > 1: return False # rapidfuzz返回0-100的匹配分数,转换为0-1的比例 ratio = fuzz.ratio(marking, code) / 100 return ratio == 1.0 or (0.98 <= ratio < 0.99)
内容的提问来源于stack exchange,提问作者dimas setiawan
相关产品推荐
相关产品推荐

