You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中编写re.findall正则提取YouTube时间戳对应描述

解决方案

现有代码问题

  • 正则仅支持MM:SS格式时间戳,未覆盖长视频常见的HH:MM:SS格式
  • 分开匹配两种时间戳位置的逻辑会产生重复数据,无法统一去重
  • 未按行匹配,可能抓取到跨行的无效内容
  • 导入pandas读取csv后未使用对应变量,属于冗余代码

优化后实现

直接替换你的get_links函数即可适配所有时间戳位置的场景:

def get_links(description):
    # 适配HH:MM:SS、MM:SS两种时间戳格式,不管时间戳在行首还是行尾都能匹配
    timestamp_pattern = re.compile(r'^.*?(?:\d{1,2}:)?\d{1,2}:\d{1,2}.*?$', re.MULTILINE)
    # 匹配所有含时间戳的行,去除每行首尾空白后返回
    timestamp_lines = [line.strip() for line in timestamp_pattern.findall(description) if line.strip()]
    
    # 输出结果,也可以根据需求返回该列表
    for line in timestamp_lines:
        print(line)
    print()

可选扩展:拆分时间戳和对应文本

如果需要把时间戳和描述文本单独拆分存储,可以用下面的逻辑:

def get_links(description):
    timestamp_lines = []
    # 匹配行内任意位置的时间戳,同时捕获时间戳和前后文本
    split_pattern = re.compile(r'(.*?)(\d{1,2}:\d{2}(?::\d{2})?)(.*)', re.MULTILINE)
    for match in split_pattern.finditer(description):
        pre_text, timestamp, post_text = match.groups()
        # 合并文本,去除多余空白
        full_text = f"{pre_text.strip()} {post_text.strip()}".strip()
        timestamp_lines.append({
            "timestamp": timestamp,
            "content": full_text
        })
    
    # 打印结果示例
    for item in timestamp_lines:
        print(f"时间戳:{item['timestamp']},描述:{item['content']}")
    print()

内容的提问来源于stack exchange,提问作者MarkWP

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.05 02:21:01