如何替换Spacy默认SentenceSegmenter为自定义分句器?(Spacy 3.4.1)
解决Spacy自定义分句器不生效的问题
问题原因
你添加的自定义组件在parser之前执行,但en_core_web_sm的parser组件会基于训练好的模型重新预测句子边界,覆盖你设置的is_sent_start标记,导致自定义规则失效。
解决方案
方案1:禁用默认parser,使用自定义分句逻辑
如果不需要句法分析功能,直接移除parser组件,只保留自定义分句逻辑:
import spacy from spacy.language import Language nlp = spacy.load("en_core_web_sm", disable=["parser"]) # 禁用parser组件 @Language.component("custom_sentencizer") def custom_sentencizer(doc): # 初始化所有token的is_sent_start为False for token in doc: token.is_sent_start = False # 设置第一个token为句子开头 doc[0].is_sent_start = True # 根据换行符设置句子开头,避免索引越界 for token in doc: if token.text == "\n" and token.i + 1 < len(doc): doc[token.i + 1].is_sent_start = True return doc # 添加自定义分句组件 nlp.add_pipe("custom_sentencizer") mystring = nlp(u"This is a sentence. This is another.\n\nThis is a\nthird sentence.") for sent in mystring.sents: print(sent)
方案2:保留parser,在其后覆盖分句结果
如果需要保留句法分析功能,将自定义组件放在parser之后运行,覆盖它的分句结果:
import spacy from spacy.language import Language nlp = spacy.load("en_core_web_sm") @Language.component("custom_sentencizer") def custom_sentencizer(doc): # 重置所有句子开头标记 for token in doc: token.is_sent_start = False doc[0].is_sent_start = True # 根据换行符设置句子开头 for token in doc: if token.text == "\n" and token.i + 1 < len(doc): doc[token.i + 1].is_sent_start = True return doc # 将组件添加到parser之后 nlp.add_pipe("custom_sentencizer", after="parser") mystring = nlp(u"This is a sentence. This is another.\n\nThis is a\nthird sentence.") for sent in mystring.sents: print(sent)
方案3:基于Spacy原生Sentencizer自定义规则
继承Spacy内置的Sentencizer类,扩展分句规则,将换行符纳入分句标记:
import spacy from spacy.pipeline.sentencizer import Sentencizer from spacy.language import Language @Language.factory("custom_sentencizer") def create_custom_sentencizer(nlp, name): # 在默认标点分句的基础上,添加换行符作为分句标记 rules = {"punct_chars": Sentencizer.default_punct_chars + ["\n"]} return Sentencizer(nlp.vocab, **rules) nlp = spacy.load("en_core_web_sm", disable=["parser"]) nlp.add_pipe("custom_sentencizer") mystring = nlp(u"This is a sentence. This is another.\n\nThis is a\nthird sentence.") for sent in mystring.sents: print(sent)
最终输出
运行上述任意方案代码,都会得到你期望的结果:
This is a sentence. This is another. This is a third sentence.
内容的提问来源于stack exchange,提问作者Crusader
相关产品推荐
相关产品推荐

