You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Apple Holdings与Washington D.C.各作为单个token进行分词?

针对spaCy和NLTK合并多词命名实体为单个Token的解决方案

spaCy 实现方式

方法1:使用PhraseMatcher批量合并指定实体

这是最直接高效的方式,适合批量处理已知的多词实体:

import spacy
from spacy.matcher import PhraseMatcher

# 加载spaCy模型
nlp = spacy.load("en_core_web_sm")
# 初始化PhraseMatcher
matcher = PhraseMatcher(nlp.vocab)

# 定义需要合并的多词实体列表
target_phrases = ["Apple Holdings", "Washington D.C."]
# 将短语转换为spaCy Doc对象
phrase_docs = [nlp(text) for text in target_phrases]
# 添加到匹配器
matcher.add("MULTI_WORD_ENTITIES", None, *phrase_docs)

# 处理示例文本
text = "Apple Holdings announced a new plan in Washington D.C. today."
doc = nlp(text)

# 遍历匹配结果,合并tokens
with doc.retokenize() as retokenizer:
    for match_id, start, end in matcher(doc):
        retokenizer.merge(doc[start:end])

# 输出处理后的tokens
for token in doc:
    print(token.text)

方法2:自定义正则匹配规则(处理带特殊格式的实体)

针对像Washington D.C.这类含标点的实体,可通过调整分词正则避免拆分:

import spacy

nlp = spacy.load("en_core_web_sm")

# 自定义分词规则,匹配带点的缩写结构
def custom_tokenizer(nlp):
    tokenizer = nlp.tokenizer
    # 修改中缀正则,避免拆分D.C.这类内部带点的缩写
    infix_re = spacy.util.compile_infix_regex([r'(?<=[A-Z])\.(?=[A-Z])'])
    
    return spacy.tokenizer.Tokenizer(
        nlp.vocab,
        prefix_search=tokenizer.prefix_search,
        suffix_search=tokenizer.suffix_search,
        infix_finditer=infix_re.finditer,
        token_match=None
    )

# 替换默认分词器
nlp.tokenizer = custom_tokenizer(nlp)

# 测试
doc = nlp("Washington D.C. is the capital.")
for token in doc:
    print(token.text)

NLTK 实现方式

使用MWETokenizer(多词表达式分词器)

NLTK的MWETokenizer专门用于合并指定的多词短语:

from nltk.tokenize import MWETokenizer, word_tokenize

# 初始化MWETokenizer,用元组指定需要合并的短语
mwe_tokenizer = MWETokenizer([('Apple', 'Holdings'), ('Washington', 'D.C.')])

# 先基础分词,再合并多词实体
text = "Apple Holdings announced a new plan in Washington D.C. today."
tokens = word_tokenize(text)
merged_tokens = mwe_tokenizer.tokenize(tokens)

print(merged_tokens)

通用建议

  • 维护一个术语表文件(如txt、csv),批量读取需要合并的实体,避免硬编码,方便后续扩展。
  • 对带特殊符号(点、连字符等)的实体,优先用正则调整分词逻辑,比单纯短语匹配更灵活。
  • 若需同时识别实体类型,spaCy可在合并token后手动添加实体标签,或结合内置NER模型使用。

内容的提问来源于stack exchange,提问作者ray

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.17 02:10:52