如何在Pandas中为拆分多行的字幕按句尾标点生成句子编号
实现思路
1. 基础句子编号实现
你已经有了标记句尾句号的presence_of_period列,用cumsum累加就能直接生成句子编号,逻辑是:每出现一次句尾句号,下一行开始就属于新的句子。
代码示例:
import pandas as pd # 构造示例数据 values=[ ['This is an example of subtitle.'], ['I want to group by sentences, which'], ['the end is determined by a period.'], ['row 0 should have sentece_1, rows 1 and 2 '], ['should have sentence_2. and this'], ['last row should have sentence_3.'] ] df=pd.DataFrame(values,columns=['subtitle']) # 标记是否存在句尾句号 df['presence_of_period'] = df['subtitle'].str.contains('\.') # 生成句子编号 df['sentence_number'] = 'sentence_' + (df['presence_of_period'].shift(fill_value=False).cumsum() + 1).astype(str)
运行后得到的sentence_number列和你示例中的需求完全一致。
2. 处理单行包含两句内容的情况
你提到的第4行"and this"需要移到下一行开头的需求,可以通过拆分单行多句内容解决,不需要单独提问,处理逻辑如下:
- 用正则按句号后的位置拆分每行文本,最多拆成两段(保留句号在第一段末尾)
- 把拆分后的内容拆成多行,过滤掉空内容
- 重新计算句子编号即可
代码示例:
# 拆分单行多句内容,按句号后位置拆分,最多拆2段 df['subtitle_split'] = df['subtitle'].str.split(r'(?<=\.)', n=1) # 爆炸成多行 df_exploded = df.explode('subtitle_split', ignore_index=True) # 去除首尾空白,过滤空行 df_exploded['subtitle_split'] = df_exploded['subtitle_split'].str.strip() df_exploded = df_exploded[df_exploded['subtitle_split'] != ''].reset_index(drop=True) # 重新计算标记和编号 df_exploded['presence_of_period'] = df_exploded['subtitle_split'].str.contains('\.') df_exploded['sentence_number'] = 'sentence_' + (df_exploded['presence_of_period'].shift(fill_value=False).cumsum() + 1).astype(str) # 重命名列得到最终结果 df_final = df_exploded.rename(columns={'subtitle_split':'subtitle'})[['subtitle', 'sentence_number', 'presence_of_period']]
运行后你会看到原来的第4行拆分后,"and this"已经单独成为新行,归到下一个句子的编号里。
内容的提问来源于stack exchange,提问作者chulo
相关产品推荐
相关产品推荐

