使用SpaCy Matcher逐行处理文本时,如何获取匹配句的上一句?
解决方案
核心思路
要获取匹配句子的上一句,关键是先将当前行的所有句子转换为可通过索引访问的列表,这样就能通过句子的位置索引直接定位到上一句。同时需要修正原代码中span被覆盖、Matcher重复初始化的问题。
修改后的代码
import spacy from spacy.matcher import Matcher import pandas as pd nlp = spacy.load("zh_core_web_sm") # 替换为你使用的模型,如en_core_web_sm pattern = [{"TEXT": "你的匹配模式"}] # 替换为你的实际匹配规则 # 初始化Matcher,移到循环外避免重复创建 matcher = Matcher(nlp.vocab) matcher.add("pattern_of_interest", [pattern]) instances = [] with open('file.txt', 'r') as f: for line in iter(f.readline, ''): line = line.strip() if not line: continue # 跳过空行 doc = nlp(line) sentences = list(doc.sents) # 将句子转换为列表,方便按索引访问 # 先收集当前行所有匹配的span及其所在句子 for match_id, start, end in matcher(doc): span = doc[start:end] # 找到span所在的句子 match_sent = None match_sent_idx = -1 for idx, sent in enumerate(sentences): if span.start >= sent.start and span.end <= sent.end: match_sent = sent match_sent_idx = idx break if match_sent: # 获取上一句:索引大于0时取前一个句子,否则设为空字符串 prev_sent = sentences[match_sent_idx - 1].text if match_sent_idx > 0 else "" instances.append(pd.Series({ "instance": str(span.text), "sentence": str(match_sent.text), "previous_sentence": prev_sent }))
关键改动说明
Matcher初始化移到循环外:原代码每次读一行都重新创建Matcher,会降低效率,提前初始化一次即可。- 句子转为列表:通过
list(doc.sents)将当前行的句子转为列表,这样可以通过索引match_sent_idx - 1直接获取上一句。 - 定位匹配span所在的句子:遍历句子列表,通过span的起始/结束位置判断它属于哪个句子,同时记录该句子的索引。
- 处理无前置句子的情况:如果匹配句子是当前行的第一句(索引为0),则上一句设为空字符串,避免索引越界。
- 避免
span被覆盖:原代码中多个匹配时,最后一个span会覆盖之前的,现在每个匹配单独处理对应的span和句子,确保数据准确。
内容的提问来源于stack exchange,提问作者microcastle
相关产品推荐
相关产品推荐

