You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何替换Spacy默认SentenceSegmenter为自定义分句器?(Spacy 3.4.1)

解决Spacy自定义分句器不生效的问题

问题原因

你添加的自定义组件在parser之前执行,但en_core_web_sm的parser组件会基于训练好的模型重新预测句子边界,覆盖你设置的is_sent_start标记,导致自定义规则失效。

解决方案

方案1:禁用默认parser,使用自定义分句逻辑

如果不需要句法分析功能,直接移除parser组件,只保留自定义分句逻辑:

import spacy
from spacy.language import Language

nlp = spacy.load("en_core_web_sm", disable=["parser"])  # 禁用parser组件

@Language.component("custom_sentencizer")
def custom_sentencizer(doc):
    # 初始化所有token的is_sent_start为False
    for token in doc:
        token.is_sent_start = False
    # 设置第一个token为句子开头
    doc[0].is_sent_start = True
    # 根据换行符设置句子开头,避免索引越界
    for token in doc:
        if token.text == "\n" and token.i + 1 < len(doc):
            doc[token.i + 1].is_sent_start = True
    return doc

# 添加自定义分句组件
nlp.add_pipe("custom_sentencizer")

mystring = nlp(u"This is a sentence. This is another.\n\nThis is a\nthird sentence.")

for sent in mystring.sents:
    print(sent)

方案2:保留parser,在其后覆盖分句结果

如果需要保留句法分析功能,将自定义组件放在parser之后运行,覆盖它的分句结果:

import spacy
from spacy.language import Language

nlp = spacy.load("en_core_web_sm")

@Language.component("custom_sentencizer")
def custom_sentencizer(doc):
    # 重置所有句子开头标记
    for token in doc:
        token.is_sent_start = False
    doc[0].is_sent_start = True
    # 根据换行符设置句子开头
    for token in doc:
        if token.text == "\n" and token.i + 1 < len(doc):
            doc[token.i + 1].is_sent_start = True
    return doc

# 将组件添加到parser之后
nlp.add_pipe("custom_sentencizer", after="parser")

mystring = nlp(u"This is a sentence. This is another.\n\nThis is a\nthird sentence.")

for sent in mystring.sents:
    print(sent)

方案3:基于Spacy原生Sentencizer自定义规则

继承Spacy内置的Sentencizer类,扩展分句规则,将换行符纳入分句标记:

import spacy
from spacy.pipeline.sentencizer import Sentencizer
from spacy.language import Language

@Language.factory("custom_sentencizer")
def create_custom_sentencizer(nlp, name):
    # 在默认标点分句的基础上,添加换行符作为分句标记
    rules = {"punct_chars": Sentencizer.default_punct_chars + ["\n"]}
    return Sentencizer(nlp.vocab, **rules)

nlp = spacy.load("en_core_web_sm", disable=["parser"])
nlp.add_pipe("custom_sentencizer")

mystring = nlp(u"This is a sentence. This is another.\n\nThis is a\nthird sentence.")

for sent in mystring.sents:
    print(sent)

最终输出

运行上述任意方案代码,都会得到你期望的结果:

This is a sentence.
This is another.

This is a
third sentence.

内容的提问来源于stack exchange,提问作者Crusader

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 10:48:22