You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

500万条语句与100个关键词的匹配性能优化技术问询

大规模文本关键词匹配性能优化方案

原代码的核心性能瓶颈在于:循环100个关键词,每个关键词都对500万条语句做全量str.contains匹配,相当于重复扫描数据集100次。以下是几种针对性的优化方案:


优化方案1:反向匹配——遍历一次数据集匹配所有关键词

不再循环关键词,改为遍历每条语句,一次性检测它包含的所有关键词,再将语句映射到对应关键词的结果列表。仅需扫描数据集1次,大幅减少IO和计算量。

from collections import defaultdict
import re

# 初始化关键词到语句列表的映射
keyword_to_sentences = defaultdict(list)
# 预编译正则表达式,转义关键词避免特殊字符干扰,按需添加\b实现精确词匹配
pattern = re.compile('|'.join([re.escape(w) for w in keywords]))
# 如需精确匹配(如"apple"不匹配"apples"),改用下方正则:
# pattern = re.compile('|'.join([r'\b' + re.escape(w) + r'\b' for w in keywords]))

# 遍历所有语句,一次匹配所有关键词
for sentence in df['sentence'].tolist():
    matched_keywords = pattern.findall(sentence)
    # 去重语句中重复出现的关键词
    unique_matches = set(matched_keywords)
    for kw in unique_matches:
        keyword_to_sentences[kw].append(sentence)

# 转换为原代码要求的context格式(按输入关键词顺序返回)
context = [keyword_to_sentences.get(kw, []) for kw in keywords]

优化方案2:Pandas矢量化批量匹配

利用Pandas底层的C优化,通过str.findall一次性提取所有语句的匹配关键词,再通过分组聚合得到结果,避免Python层面的循环开销。

import re
import pandas as pd

# 预编译正则
pattern = re.compile('|'.join([re.escape(w) for w in keywords]))
# 提取每条语句的匹配关键词列表
df['matched_kw'] = df['sentence'].str.findall(pattern)
# 展开列表并过滤无匹配的语句
exploded_df = df.explode('matched_kw').dropna(subset=['matched_kw'])
# 按关键词分组,收集对应语句
grouped = exploded_df.groupby('matched_kw')['sentence'].apply(list)
# 生成目标context列表
context = [grouped.get(kw, []) for kw in keywords]

优化方案3:构建倒排索引(适合多次查询场景)

如果后续需要多次查询不同关键词,使用专业文本检索库构建倒排索引,查询速度会远快于全量扫描。以Whoosh为例:

from whoosh.index import create_in
from whoosh.fields import Schema, TEXT
from whoosh.qparser import QueryParser
import os

# 创建索引结构
schema = Schema(sentence=TEXT(stored=True))
# 创建临时索引目录
if not os.path.exists('indexdir'):
    os.mkdir('indexdir')
ix = create_in('indexdir', schema)

# 将所有语句写入索引
writer = ix.writer()
for sentence in df['sentence'].tolist():
    writer.add_document(sentence=sentence)
writer.commit()

# 批量查询关键词
context = []
with ix.searcher() as searcher:
    parser = QueryParser("sentence", ix.schema)
    for kw in keywords:
        query = parser.parse(kw)
        results = searcher.search(query, limit=None)
        context.append([hit['sentence'] for hit in results])

优化方案4:多进程并行处理

如果机器CPU核心充足,可将数据集分片后并行处理,进一步压缩耗时(注意:小数据集可能因进程通信开销导致变慢,适合超大规模场景)。

from multiprocessing import Pool
from collections import defaultdict
import re

def process_chunk(chunk):
    chunk_result = defaultdict(list)
    pattern = re.compile('|'.join([re.escape(w) for w in keywords]))
    for sentence in chunk:
        matched = set(pattern.findall(sentence))
        for kw in matched:
            chunk_result[kw].append(sentence)
    return chunk_result

# 根据CPU核心数分片示例(此处设为4)
chunks = [df['sentence'].tolist()[i::4] for i in range(4)]

# 并行处理分片
with Pool(4) as p:
    results = p.map(process_chunk, chunks)

# 合并分片结果
keyword_to_sentences = defaultdict(list)
for res in results:
    for kw, sentences in res.items():
        keyword_to_sentences[kw].extend(sentences)

context = [keyword_to_sentences.get(kw, []) for kw in keywords]

方案选择建议

  • 一次性匹配需求:优先选反向匹配Python循环或Pandas矢量化方案,实现简单且性能提升明显;
  • 多次查询需求:优先构建倒排索引,后续查询几乎无开销;
  • 超大规模数据集+多核CPU:结合多进程并行进一步提速。

内容的提问来源于stack exchange,提问作者Guy Barash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.23 15:39:21