You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas中explode字符串列时如何保留其他列 优先使用spacy实现分句

实现代码
import pandas as pd
import spacy

# 加载英文spacy模型,处理中文请替换为zh_core_web_sm
nlp = spacy.load("en_core_web_sm")

# 示例数据
df = pd.DataFrame({'col1' : [1,2,3],
                   'col2' : ['A','B','C'],
                   'paragraph': ['sentence one. sentence two',
                                 'sentence three. and sentence four',
                                 'crazy sentence!! and the final one.']})

# 自定义分句函数
def split_sentences(text):
    doc = nlp(text)
    # 提取句子文本,清理首尾空格,过滤空字符串
    return [sent.text.strip() for sent in doc.sents if sent.text.strip()]

# 应用分句函数后拆分,自动保留其余列信息
df_exploded = df.assign(paragraph=df['paragraph'].apply(split_sentences)).explode('paragraph').reset_index(drop=True)

print(df_exploded)

输出结果:

col1 col2               paragraph
0     1    A            sentence one
1     1    A            sentence two
2     2    B          sentence three
3     2    B       and sentence four
4     3    C       crazy sentence!!
5     3    C    and the final one.
关键逻辑说明
  • 之前仅对paragraph单列做拆分操作所以会丢失其余列,改为用assign对paragraph列做处理后再调用explode,pandas会自动对齐其余列的原值,每拆分出一个句子,对应的col1、col2值都会复制保留
  • spacy的sents内置了多标点(感叹号、问号、分号等)的分句规则,无需手动处理拆分规则,适配多场景分句需求
  • 分句函数内增加了空字符串过滤和首尾空格清理,避免出现拆分后空行的问题

内容的提问来源于stack exchange,提问作者ℕʘʘḆḽḘ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 10:42:02