You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python的spacy库实现基于双换行符的段落边界检测

基于spaCy实现仅按双换行符分割段落的方案

方法一:自定义spaCy管道组件(可后续对段落做NLP标注)

如果需要在分割段落的同时,复用spaCy的其他NLP功能(比如词性标注、实体识别等),可以自定义分割组件插入spaCy的处理管道:

import spacy
from spacy.language import Language

# 注册自定义段落分割逻辑,仅在\n\n处分割
@Language.component("custom_para_splitter")
def set_custom_segmentation(doc):
    # 先默认所有位置都不允许分割
    for token in doc:
        token.is_sent_start = False
    # 仅在双换行符的下一个位置设置为分割起点
    for i in range(len(doc)-1):
        if doc[i].text == "\n" and doc[i+1].text == "\n":
            doc[i+1].is_sent_start = True
    return doc

# 加载空白英文模型,不需要预训练模型也能实现分割,节省资源
nlp = spacy.blank("en")
# 将自定义分割组件加入管道,放在最前面
nlp.add_pipe("custom_para_splitter", first=True)

# 你的原始文本
para = '''Machine learning is a branch of artificial intelligence (AI) and computer science. It focuses on the use of data and algorithms to imitate the way that humans learn, gradually improving its accuracy.

Machine learning is an important component of the growing field of data science. Through the use of statistical methods,algorithms are trained to make classifications or predictions, uncovering key insights within data mining projects.These insights subsequently drive decision making within applications and businesses, ideally impacting key growth metrics.As big data continues to expand and grow, the market demand for data scientists will increase, requiring them to assist in theidentification of the most relevant business questions and subsequently the data to answer them.'''

doc = nlp(para)
# 按分割后的句子提取,就是你要的段落
para_split = [sent.text.rstrip("\n") for sent in doc.sents]
print(para_split)
print(len(para_split))

运行后输出结果和你之前NLTK实现的完全一致:

['Machine learning is a branch of artificial intelligence (AI) and computer science. It focuses on the use of data and algorithms to imitate the way that humans learn, gradually improving its accuracy.', 'Machine learning is an important component of the growing field of data science. Through the use of statistical methods,algorithms are trained to make classifications or predictions, uncovering key insights within data mining projects.These insights subsequently drive decision making within applications and businesses, ideally impacting key growth metrics.As big data continues to expand and grow, the market demand for data scientists will increase, requiring them to assist in theidentification of the most relevant business questions and subsequently the data to answer them.']
2

方法二:轻量实现(仅需分割不需要后续NLP处理)

如果你不需要用到spaCy的其他NLP功能,只是需要匹配需求的分割结果,也可以用更简单的写法:

import spacy

nlp = spacy.blank("en")
para = "你的原始文本"
# 直接按双换行分割
para_split = para.split("\n\n")
# 如果需要把每个段落转成spaCy的Doc对象做后续处理,再加一行代码即可:
docs = [nlp(p) for p in para_split]

这种写法运行效率更高,和你需求的匹配度也完全一致。

内容的提问来源于stack exchange,提问作者Nishant

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 16:36:01