pandas中explode字符串列时如何保留其他列 优先使用spacy实现分句
实现代码
import pandas as pd import spacy # 加载英文spacy模型,处理中文请替换为zh_core_web_sm nlp = spacy.load("en_core_web_sm") # 示例数据 df = pd.DataFrame({'col1' : [1,2,3], 'col2' : ['A','B','C'], 'paragraph': ['sentence one. sentence two', 'sentence three. and sentence four', 'crazy sentence!! and the final one.']}) # 自定义分句函数 def split_sentences(text): doc = nlp(text) # 提取句子文本,清理首尾空格,过滤空字符串 return [sent.text.strip() for sent in doc.sents if sent.text.strip()] # 应用分句函数后拆分,自动保留其余列信息 df_exploded = df.assign(paragraph=df['paragraph'].apply(split_sentences)).explode('paragraph').reset_index(drop=True) print(df_exploded)
输出结果:
col1 col2 paragraph 0 1 A sentence one 1 1 A sentence two 2 2 B sentence three 3 2 B and sentence four 4 3 C crazy sentence!! 5 3 C and the final one.
关键逻辑说明
- 之前仅对
paragraph单列做拆分操作所以会丢失其余列,改为用assign对paragraph列做处理后再调用explode,pandas会自动对齐其余列的原值,每拆分出一个句子,对应的col1、col2值都会复制保留 - spacy的
sents内置了多标点(感叹号、问号、分号等)的分句规则,无需手动处理拆分规则,适配多场景分句需求 - 分句函数内增加了空字符串过滤和首尾空格清理,避免出现拆分后空行的问题
内容的提问来源于stack exchange,提问作者ℕʘʘḆḽḘ
相关产品推荐
相关产品推荐

