spaCy正则匹配Span的语言工厂实现及跨文件复用方案
解决spaCy自定义Language Factory的[E966]错误及复用问题
1. [E966]错误原因
- spaCy的
nlp.add_pipe()要求添加已注册的工厂名称,而非直接传入组件实例(可调用对象)。如果跳过工厂注册步骤直接传组件,就会触发该错误。
2. 正确实现并注册Language Factory
步骤1:编写正则匹配Span的组件
先实现核心逻辑,处理文档文本的正则匹配并生成Span:
import re from spacy.tokens import Doc, Span from spacy.language import Language class RegexSpanMatcher: def __init__(self, nlp, pattern: str, label: str): self.nlp = nlp self.pattern = re.compile(pattern) self.label = label # 注册Span扩展属性(可选,标记是否为正则匹配结果) if not Span.has_extension("regex_match"): Span.set_extension("regex_match", default=False) def __call__(self, doc: Doc) -> Doc: matches = self.pattern.finditer(doc.text) spans = [] for match in matches: char_start, char_end = match.span() # 将字符索引转换为token索引,处理匹配结果为空的情况 token_span = doc.char_span(char_start, char_end) if token_span: span = Span(doc, token_span.start, token_span.end, label=self.label) span._.regex_match = True spans.append(span) # 将匹配的Span加入文档实体(若无需作为实体,可存储到自定义Doc扩展属性) doc.ents = list(doc.ents) + spans return doc
步骤2:注册为Language Factory
使用@Language.factory装饰器将组件注册为可复用的工厂:
@Language.factory( "regex_span_matcher", # 工厂名称,后续add_pipe时需使用此名称 default_config={"pattern": r"\b\d{3}-\d{2}-\d{4}\b", "label": "SSN"} # 默认参数配置 ) def create_regex_span_matcher(nlp, name, pattern, label): # 工厂函数负责初始化组件实例 return RegexSpanMatcher(nlp, pattern, label)
步骤3:调用nlp.add_pipe(避免[E966]错误)
现在通过工厂名称添加组件,即可正常运行:
import spacy nlp = spacy.load("en_core_web_sm") # 可通过config参数覆盖默认的pattern和label nlp.add_pipe("regex_span_matcher", config={"pattern": r"\b[A-Z]{2}\d{4}\b", "label": "USER_ID"}) # 测试效果 doc = nlp("My user ID is XY7890 and SSN is 123-45-6789") for ent in doc.ents: print(f"{ent.text}: {ent.label_}")
3. 保存工厂到.py文件并复用
步骤1:封装到独立文件
创建custom_spacy_components.py,将上述RegexSpanMatcher类和工厂注册代码全部存入该文件。
步骤2:在其他文件中调用
只需导入该文件,spaCy会自动加载已注册的工厂,直接通过名称使用:
import spacy # 导入自定义组件文件,触发工厂注册 from custom_spacy_components import RegexSpanMatcher nlp = spacy.load("en_core_web_sm") # 添加组件并自定义参数 nlp.add_pipe("regex_span_matcher", config={"pattern": r"\b\d{11}\b", "label": "PHONE"}) # 测试 doc = nlp("My phone number is 13800138000") for ent in doc.ents: print(f"{ent.text}: {ent.label_}")
额外提示
- 若需将包含自定义组件的管道保存到磁盘(
nlp.to_disk()),需为组件实现to_disk和from_disk方法,确保参数能正确序列化 - 处理
char_span返回None的情况(如匹配文本跨token边界且无法映射到完整token),可添加日志或跳过无效匹配
内容的提问来源于stack exchange,提问作者JFerro
相关产品推荐
相关产品推荐

