You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中提取字符串第三个时间标记及后续内容?

处理Pandas列中带时间标记的字符串

没问题,这事儿用Pandas结合正则就能轻松搞定,给你写个直接能用的方案:

import pandas as pd
import re

# 先整个示例数据方便你测试
df = pd.DataFrame({
    'text': [
        "(15:38) Hello, how are you? (15:39) I am fine. (15:40) That's good. ",
        "(09:01) Hi there! (09:02) Bye!",
        "(12:30) Single message"
    ]
})

def get_third_and_beyond(text):
    # 先把所有时间标记提取出来
    time_tags = re.findall(r'\(\d{1,2}:\d{1,2}\)', text)
    # 再按时间标记分割内容
    content_parts = re.split(r'\(\d{1,2}:\d{1,2}\)', text)
    
    # 检查时间标记够不够3个
    if len(time_tags) >= 3:
        # 从第三个标记开始,把标记和对应的内容拼回去
        result = ''.join([f"{tag}{part}" for tag, part in zip(time_tags[2:], content_parts[3:])])
        return result.strip()
    # 不够的话直接返回空字符串
    else:
        return ''

# 把函数应用到你的目标列上
df['processed_text'] = df['text'].apply(get_third_and_beyond)

运行后你看df['processed_text']的结果,完全符合你的要求:

0    (15:40) That's good.
1                        
2                        
Name: processed_text, dtype: object

为啥这么写?

  • 用re.findall()把所有时间标记先捞出来,这样能直接数清楚数量,方便判断是否满足≥3的条件。
  • 用re.split()分割后,第一个元素是空字符串(因为原字符串开头就是时间标记),所以content_parts的长度比时间标记数多1,zip(time_tags[2:], content_parts[3:])刚好能把第三个及以后的标记和对应的内容一一配对,拼回去就是你要的内容。
  • 最后加个strip()是为了去掉首尾多余的空格,看起来更干净。

内容的提问来源于stack exchange,提问作者Dylan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:37:21