You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Spacy Matcher实现含排除项的松散单词匹配?

解决Spacy Matcher匹配排除特定token的问题

当然可以实现这个需求!你的问题出在当前模式的逻辑上:{'OP': '*'}会匹配任意数量(包括0)的任意token,包括你想排除的"store",而后面的{'OP': '!', 'LOWER': 'store'}并没有起到预期的排除作用——它只是要求在那个特定位置没有"store",但前面的*已经把"store"包含在匹配序列里了,所以"play store game"还是会被匹配到。

修改后的代码

from spacy.matcher import Matcher
import spacy

nlp = spacy.load("en_core_web_sm")  # 补充原代码缺失的模型加载步骤
matcher = Matcher(nlp.vocab, validate=True)

# 调整模式:确保play和目标词之间的所有token都不包含store
pattern = [
    {'LOWER': 'play'},
    {'LOWER': {'NOT_IN': ['store']}, 'OP': '*'},  # 匹配任意数量非"store"的token
    {'LOWER': {'IN': ["game", "pacman"]}}
]
matcher.add('HUNTING', None, pattern)

def extract_patterns(nlp_doc, matcher):
    result_spans = []
    matches = matcher(nlp_doc)
    print("matches:", len(matches))
    for match_id, start, end in matches:
        span = nlp_doc[start:end]
        result_spans.append(span)
    return result_spans

text = ('play store game. \n play with pacman')
doc = nlp(text)
output = extract_patterns(doc, matcher=matcher)
print([str(span) for span in output])  # 输出: ['play with pacman']

关键逻辑说明

  • {'LOWER': {'NOT_IN': ['store']}, 'OP': '*'}:这部分规则表示可以匹配零个或多个token,但每个token的小写形式都不能是"store",从根源上排除了中间包含"store"的序列。
  • 原代码遗漏了nlp模型加载步骤,我已经补充完整,否则代码无法正常运行。

额外补充

如果后续需要更复杂的排除逻辑(比如排除包含"store"的完整短语),可以考虑结合PhraseMatcher或者在匹配后添加自定义过滤逻辑,但针对你当前的需求,上面的模式调整已经完全满足要求。

内容的提问来源于stack exchange,提问作者Pamin Rangsikunpum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:57:17