如何使用Spacy Matcher实现含排除项的松散单词匹配?
解决Spacy Matcher匹配排除特定token的问题
当然可以实现这个需求!你的问题出在当前模式的逻辑上:{'OP': '*'}会匹配任意数量(包括0)的任意token,包括你想排除的"store",而后面的{'OP': '!', 'LOWER': 'store'}并没有起到预期的排除作用——它只是要求在那个特定位置没有"store",但前面的*已经把"store"包含在匹配序列里了,所以"play store game"还是会被匹配到。
修改后的代码
from spacy.matcher import Matcher import spacy nlp = spacy.load("en_core_web_sm") # 补充原代码缺失的模型加载步骤 matcher = Matcher(nlp.vocab, validate=True) # 调整模式:确保play和目标词之间的所有token都不包含store pattern = [ {'LOWER': 'play'}, {'LOWER': {'NOT_IN': ['store']}, 'OP': '*'}, # 匹配任意数量非"store"的token {'LOWER': {'IN': ["game", "pacman"]}} ] matcher.add('HUNTING', None, pattern) def extract_patterns(nlp_doc, matcher): result_spans = [] matches = matcher(nlp_doc) print("matches:", len(matches)) for match_id, start, end in matches: span = nlp_doc[start:end] result_spans.append(span) return result_spans text = ('play store game. \n play with pacman') doc = nlp(text) output = extract_patterns(doc, matcher=matcher) print([str(span) for span in output]) # 输出: ['play with pacman']
关键逻辑说明
{'LOWER': {'NOT_IN': ['store']}, 'OP': '*'}:这部分规则表示可以匹配零个或多个token,但每个token的小写形式都不能是"store",从根源上排除了中间包含"store"的序列。- 原代码遗漏了
nlp模型加载步骤,我已经补充完整,否则代码无法正常运行。
额外补充
如果后续需要更复杂的排除逻辑(比如排除包含"store"的完整短语),可以考虑结合PhraseMatcher或者在匹配后添加自定义过滤逻辑,但针对你当前的需求,上面的模式调整已经完全满足要求。
内容的提问来源于stack exchange,提问作者Pamin Rangsikunpum
相关产品推荐
相关产品推荐

