You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中为拆分多行的字幕按句尾标点生成句子编号

实现思路

1. 基础句子编号实现

你已经有了标记句尾句号的presence_of_period列,用cumsum累加就能直接生成句子编号,逻辑是:每出现一次句尾句号,下一行开始就属于新的句子。
代码示例:

import pandas as pd

# 构造示例数据
values=[
        ['This is an example of subtitle.'],
        ['I want to group by sentences, which'],
        ['the end is determined by a period.'],
        ['row 0 should have sentece_1, rows 1 and 2 '],
        ['should have sentence_2. and this'],
        ['last row should have sentence_3.']
        ]
df=pd.DataFrame(values,columns=['subtitle'])

# 标记是否存在句尾句号
df['presence_of_period'] = df['subtitle'].str.contains('\.')

# 生成句子编号
df['sentence_number'] = 'sentence_' + (df['presence_of_period'].shift(fill_value=False).cumsum() + 1).astype(str)

运行后得到的sentence_number列和你示例中的需求完全一致。


2. 处理单行包含两句内容的情况

你提到的第4行"and this"需要移到下一行开头的需求,可以通过拆分单行多句内容解决,不需要单独提问,处理逻辑如下:

  • 用正则按句号后的位置拆分每行文本,最多拆成两段(保留句号在第一段末尾)
  • 把拆分后的内容拆成多行,过滤掉空内容
  • 重新计算句子编号即可
    代码示例:
# 拆分单行多句内容,按句号后位置拆分,最多拆2段
df['subtitle_split'] = df['subtitle'].str.split(r'(?<=\.)', n=1)
# 爆炸成多行
df_exploded = df.explode('subtitle_split', ignore_index=True)
# 去除首尾空白,过滤空行
df_exploded['subtitle_split'] = df_exploded['subtitle_split'].str.strip()
df_exploded = df_exploded[df_exploded['subtitle_split'] != ''].reset_index(drop=True)

# 重新计算标记和编号
df_exploded['presence_of_period'] = df_exploded['subtitle_split'].str.contains('\.')
df_exploded['sentence_number'] = 'sentence_' + (df_exploded['presence_of_period'].shift(fill_value=False).cumsum() + 1).astype(str)

# 重命名列得到最终结果
df_final = df_exploded.rename(columns={'subtitle_split':'subtitle'})[['subtitle', 'sentence_number', 'presence_of_period']]

运行后你会看到原来的第4行拆分后,"and this"已经单独成为新行,归到下一个句子的编号里。


内容的提问来源于stack exchange,提问作者chulo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 21:00:00