You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何向spaCy添加自定义词,让Tokenizer识别多词生物名称为单个token

解决spaCy拆分多词生物拉丁名的方案

方法1:用Matcher批量处理上千个自定义术语

这是处理大规模词汇最高效的方式,步骤如下:

  1. 先把你的生物名称整理成一个列表:
custom_terms = ["Angelica sinensis", "Homo sapiens", "Panthera leo"]
# 替换成你自己的上千个术语列表
  1. 用spaCy的Matcher创建匹配规则,再通过自定义组件合并匹配到的多词序列:
import spacy
from spacy.matcher import Matcher

# 医学场景推荐用sciSpacy的模型,对生物术语适配性更强
nlp = spacy.load("en_core_sci_sm")
matcher = Matcher(nlp.vocab)

# 为每个术语生成匹配模式(自动忽略大小写差异)
for term in custom_terms:
    pattern = [{"LOWER": word.lower()} for word in term.split()]
    matcher.add(term, [pattern])

# 定义合并匹配术语的组件
def merge_terms(doc):
    matches = matcher(doc)
    spans = []
    for match_id, start, end in matches:
        spans.append(doc[start:end])
    # 过滤重叠的匹配项,避免重复合并
    for span in spacy.util.filter_spans(spans):
        span.merge()
    return doc

# 把组件加到管道里,放在tokenizer之后执行
nlp.add_pipe(merge_terms, after="tokenizer")

# 测试效果
doc = nlp("Angelica sinensis is a widely used medicinal herb.")
print([token.text for token in doc])
# 输出:['Angelica sinensis', 'is', 'a', 'widely', 'used', 'medicinal', 'herb', '.']

方法2:给Tokenizer加特殊规则(适合少量术语)

如果术语数量少可以用这种方式,但上千个的话会让规则太臃肿、效率下降,仅作参考:

nlp = spacy.load("en_core_sci_sm")
# 强制让Tokenizer把指定多词序列识别为单个token
nlp.tokenizer.add_special_case("Angelica sinensis", [{"ORTH": "Angelica sinensis"}])

doc = nlp("Angelica sinensis has therapeutic effects.")
print([token.text for token in doc])
# 输出:['Angelica sinensis', 'has', 'therapeutic', 'effects', '.']

实用提醒

  • 优先用方法1处理上千个术语,维护成本和运行效率都更优。
  • sciSpacy的预训练模型比通用模型更适配医学文献,能减少很多基础的术语拆分问题。
  • 术语列表记得先去重,避免重复添加规则拖慢处理速度。

内容的提问来源于stack exchange,提问作者small tomato

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 22:05:16