如何在spaCy Matcher中设置匹配规则优先级?
如何在spaCy Matcher中设置规则优先级避免重复匹配
需求:对于句子"There is no apple or is there an apple?",希望no apple的匹配规则优先,当该规则匹配时,不返回单独的apple匹配结果。
尝试1:单个模式包含两种子规则
尝试用一个模式同时检测非no开头的apple和单独的apple,代码如下:
import spacy from spacy.matcher import Matcher nlp = spacy.load("en_core_web_sm") matcher = Matcher(nlp.vocab) pattern = [ [{"LOWER": {"NOT_IN": ["no"]}}, {"LOWER": "apple"}], [{"LOWER": "apple"}] ] matcher.add("apple", pattern) doc = nlp("There is no apple or is there an apple?") matches = matcher(doc) for match_id, start, end in matches: string_id = nlp.vocab.strings[match_id] span = doc[start:end] print(match_id, string_id, start, end, span.text)
输出结果:
8566208034543834098 apple 3 4 apple 8566208034543834098 apple 7 9 an apple 8566208034543834098 apple 8 9 apple
问题:第二个子模式导致多次匹配apple,且无法排除no apple中的apple。
尝试2:独立的两个匹配规则
创建no_apple和apple两个独立模式,代码如下:
import spacy from spacy.matcher import Matcher nlp = spacy.load("en_core_web_sm") matcher = Matcher(nlp.vocab) pattern = [ [{"LOWER": "apple"}], ] no_pattern = [ [{"LOWER": "no"}, {"LOWER": "apple"}], ] matcher.add("apple", pattern) matcher.add("no_apple", no_pattern) doc = nlp("There is no apple or is there an apple?") matches = matcher(doc) for match_id, start, end in matches: string_id = nlp.vocab.strings[match_id] span = doc[start:end] print(match_id, string_id, start, end, span.text)
输出结果:
14541201340755442066 no_apple 2 4 no apple 8566208034543834098 apple 3 4 apple 8566208034543834098 apple 8 9 apple
问题:no_apple和apple会同时匹配,存在重复匹配的问题。
解决方案:手动实现优先级与重叠过滤
spaCy Matcher本身没有内置的规则优先级配置,但可以通过排序+过滤重叠匹配的方式实现需求,步骤如下:
- 定义所有匹配规则,明确优先级(比如
no_apple优先级高于apple)。 - 获取所有匹配结果后,按优先级降序、匹配长度降序排序(长匹配优先)。
- 遍历排序后的结果,过滤掉与已保留匹配重叠的低优先级结果。
代码示例:
import spacy from spacy.matcher import Matcher nlp = spacy.load("en_core_web_sm") matcher = Matcher(nlp.vocab) # 添加匹配规则 no_pattern = [[{"LOWER": "no"}, {"LOWER": "apple"}]] matcher.add("no_apple", no_pattern) pattern = [[{"LOWER": "apple"}]] matcher.add("apple", pattern) doc = nlp("There is no apple or is there an apple?") matches = matcher(doc) # 定义规则优先级:no_apple > apple priority_map = {"no_apple": 2, "apple": 1} # 按优先级降序、匹配长度降序排序 sorted_matches = sorted(matches, key=lambda x: (-priority_map[nlp.vocab.strings[x[0]]], -(x[2]-x[1]))) # 过滤重叠匹配,保留高优先级结果 filtered_matches = [] used_tokens = set() for match_id, start, end in sorted_matches: token_range = set(range(start, end)) if not token_range & used_tokens: filtered_matches.append((match_id, start, end)) used_tokens.update(token_range) # 按原文本顺序输出结果 filtered_matches.sort(key=lambda x: x[1]) for match_id, start, end in filtered_matches: string_id = nlp.vocab.strings[match_id] span = doc[start:end] print(match_id, string_id, start, end, span.text)
输出结果:
14541201340755442066 no_apple 2 4 no apple 8566208034543834098 apple 8 9 apple
也可以使用on_match回调函数在匹配过程中实时过滤,适合简单场景,但排序+过滤的方式更通用,能处理复杂的优先级规则。
内容的提问来源于stack exchange,提问作者Quinten
相关产品推荐
相关产品推荐

