如何基于Document-level Inverted Index高效构建Sentence-level Inverted Index?
优化从文档级倒排索引构建句子级倒排索引的速度
看起来你现在的瓶颈主要在正则匹配的效率上——当related_word_list比较大时,用'|'.join生成的正则表达式会有大量分支,每个句子都要遍历这些分支,自然慢得离谱。下面几个优化思路可以帮你把速度提上来,不用多进程也能达到每秒处理10个文档的目标:
1. 反向思路:用集合匹配替代正则(最推荐)
与其让每个句子去匹配所有相关词的正则,不如先从文档级倒排索引里拿到当前文档的相关词集合(用set存储,查找是O(1)复杂度),然后对每个句子分词后,直接检查哪些词在这个集合里。这种方法的效率比正则高几个数量级。
代码示例:
from collections import defaultdict # 第一步:从文档级倒排索引构建「文档ID -> 相关词集合」的映射 doc_to_related_terms = {} for term, doc_ids in doc_level_inverted_index.items(): for doc_id in doc_ids: if doc_id not in doc_to_related_terms: doc_to_related_terms[doc_id] = set() doc_to_related_terms[doc_id].add(term.lower()) # 统一小写,避免大小写不匹配 # 第二步:遍历每个文档的句子,构建句子级倒排索引 sentence_level_index = defaultdict(list) for doc_id, doc in documents.items(): related_terms = doc_to_related_terms.get(doc_id, set()) if not related_terms: continue # 没有相关词的文档直接跳过 for sent_idx, sentence in enumerate(doc.sentences): # 简单分词:处理标点、转小写,也可以用更专业的分词工具 tokens = [token.strip('.,!?;:') for token in sentence.lower().split()] # 找当前句子和相关词的交集 matched_terms = related_terms.intersection(tokens) for term in matched_terms: sentence_level_index[term].append(f"{doc_id}-sent{sent_idx}")
为什么快?
set的交集操作是底层优化过的,比正则的分支匹配快得多;- 只关注当前文档的相关词,而不是全局所有词,减少了匹配范围;
- 逻辑简单,没有正则的额外开销。
2. 优化正则匹配(如果必须用正则)
如果你需要精确的单词匹配(比如避免匹配到包含目标词的更长词汇,比如aaa匹配到aaab),可以预编译针对每个文档的正则,而不是全局大正则,同时加入单词边界限制。
代码示例:
import re from collections import defaultdict # 第一步:构建「文档ID -> 相关词集合」的映射 doc_to_terms = {} for term, doc_ids in doc_level_inverted_index.items(): for doc_id in doc_ids: if doc_id not in doc_to_terms: doc_to_terms[doc_id] = set() doc_to_terms[doc_id].add(term) # 预编译每个文档的正则 doc_to_regex = {} for doc_id, terms in doc_to_terms.items(): # 用re.escape处理特殊字符,加\b确保匹配完整单词 pattern_str = r'\b(' + '|'.join(re.escape(t) for t in terms) + r')\b' doc_to_regex[doc_id] = re.compile(pattern_str, re.IGNORECASE) # 第二步:处理文档 sentence_level_index = defaultdict(list) for doc_id, doc in documents.items(): regex = doc_to_regex.get(doc_id) if not regex: continue for sent_idx, sentence in enumerate(doc.sentences): # 找到所有匹配的词 matched_terms = regex.findall(sentence) for term in matched_terms: sentence_level_index[term.lower()].append(f"{doc_id}-sent{sent_idx}")
为什么比原代码快?
- 预编译正则避免了重复编译的开销;
- 每个文档的正则只包含自己的相关词,分支更少;
\b减少了无效匹配,正则引擎的匹配效率更高。
3. 用高效分词工具提升精度和速度
如果你的分词逻辑(比如简单的split)不够准确,导致匹配错误或者额外开销,可以用专业的分词库,比如spaCy的轻量级分词,它处理标点、大小写、复合词的效率和精度都比手动split高。
代码示例:
import spacy from collections import defaultdict # 加载轻量级spaCy模型,禁用不需要的组件(只保留分词) nlp = spacy.load("en_core_web_sm", disable=["tagger", "parser", "ner"]) # 第一步:构建文档到相关词的映射(同方案1) doc_to_related_terms = {} for term, doc_ids in doc_level_inverted_index.items(): for doc_id in doc_ids: if doc_id not in doc_to_related_terms: doc_to_related_terms[doc_id] = set() doc_to_related_terms[doc_id].add(term.lower()) # 第二步:用spaCy分词处理 sentence_level_index = defaultdict(list) for doc_id, doc in documents.items(): related_terms = doc_to_related_terms.get(doc_id, set()) if not related_terms: continue for sent_idx, sentence in enumerate(doc.sentences): # spaCy快速分词 doc_spacy = nlp(sentence) tokens = [token.text.lower() for token in doc_spacy] matched_terms = related_terms.intersection(tokens) for term in matched_terms: sentence_level_index[term].append(f"{doc_id}-sent{sent_idx}")
为什么好?
- spaCy的分词是C优化过的,速度极快;
- 自动处理标点、大小写、连字符等情况,不用手动写复杂的字符串处理逻辑;
- 分词结果更准确,减少误匹配或漏匹配。
内容的提问来源于stack exchange,提问作者Tom Leung
相关产品推荐
相关产品推荐

