You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

spaCy正则匹配Span的语言工厂实现及跨文件复用方案

解决spaCy自定义Language Factory的[E966]错误及复用问题

1. [E966]错误原因

  • spaCy的nlp.add_pipe()要求添加已注册的工厂名称,而非直接传入组件实例(可调用对象)。如果跳过工厂注册步骤直接传组件,就会触发该错误。

2. 正确实现并注册Language Factory

步骤1:编写正则匹配Span的组件

先实现核心逻辑,处理文档文本的正则匹配并生成Span:

import re
from spacy.tokens import Doc, Span
from spacy.language import Language

class RegexSpanMatcher:
    def __init__(self, nlp, pattern: str, label: str):
        self.nlp = nlp
        self.pattern = re.compile(pattern)
        self.label = label
        # 注册Span扩展属性(可选,标记是否为正则匹配结果)
        if not Span.has_extension("regex_match"):
            Span.set_extension("regex_match", default=False)

    def __call__(self, doc: Doc) -> Doc:
        matches = self.pattern.finditer(doc.text)
        spans = []
        for match in matches:
            char_start, char_end = match.span()
            # 将字符索引转换为token索引,处理匹配结果为空的情况
            token_span = doc.char_span(char_start, char_end)
            if token_span:
                span = Span(doc, token_span.start, token_span.end, label=self.label)
                span._.regex_match = True
                spans.append(span)
        # 将匹配的Span加入文档实体(若无需作为实体,可存储到自定义Doc扩展属性)
        doc.ents = list(doc.ents) + spans
        return doc

步骤2:注册为Language Factory

使用@Language.factory装饰器将组件注册为可复用的工厂:

@Language.factory(
    "regex_span_matcher",  # 工厂名称,后续add_pipe时需使用此名称
    default_config={"pattern": r"\b\d{3}-\d{2}-\d{4}\b", "label": "SSN"}  # 默认参数配置
)
def create_regex_span_matcher(nlp, name, pattern, label):
    # 工厂函数负责初始化组件实例
    return RegexSpanMatcher(nlp, pattern, label)

步骤3:调用nlp.add_pipe(避免[E966]错误)

现在通过工厂名称添加组件,即可正常运行:

import spacy

nlp = spacy.load("en_core_web_sm")
# 可通过config参数覆盖默认的pattern和label
nlp.add_pipe("regex_span_matcher", config={"pattern": r"\b[A-Z]{2}\d{4}\b", "label": "USER_ID"})

# 测试效果
doc = nlp("My user ID is XY7890 and SSN is 123-45-6789")
for ent in doc.ents:
    print(f"{ent.text}: {ent.label_}")

3. 保存工厂到.py文件并复用

步骤1:封装到独立文件

创建custom_spacy_components.py,将上述RegexSpanMatcher类和工厂注册代码全部存入该文件。

步骤2:在其他文件中调用

只需导入该文件,spaCy会自动加载已注册的工厂,直接通过名称使用:

import spacy
# 导入自定义组件文件,触发工厂注册
from custom_spacy_components import RegexSpanMatcher

nlp = spacy.load("en_core_web_sm")
# 添加组件并自定义参数
nlp.add_pipe("regex_span_matcher", config={"pattern": r"\b\d{11}\b", "label": "PHONE"})

# 测试
doc = nlp("My phone number is 13800138000")
for ent in doc.ents:
    print(f"{ent.text}: {ent.label_}")

额外提示

  • 若需将包含自定义组件的管道保存到磁盘(nlp.to_disk()),需为组件实现to_disk和from_disk方法,确保参数能正确序列化
  • 处理char_span返回None的情况(如匹配文本跨token边界且无法映射到完整token),可添加日志或跳过无效匹配

内容的提问来源于stack exchange,提问作者JFerro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 14:40:46