如何在Python中编写re.findall正则提取YouTube时间戳对应描述
解决方案
现有代码问题
- 正则仅支持
MM:SS格式时间戳,未覆盖长视频常见的HH:MM:SS格式 - 分开匹配两种时间戳位置的逻辑会产生重复数据,无法统一去重
- 未按行匹配,可能抓取到跨行的无效内容
- 导入pandas读取csv后未使用对应变量,属于冗余代码
优化后实现
直接替换你的get_links函数即可适配所有时间戳位置的场景:
def get_links(description): # 适配HH:MM:SS、MM:SS两种时间戳格式,不管时间戳在行首还是行尾都能匹配 timestamp_pattern = re.compile(r'^.*?(?:\d{1,2}:)?\d{1,2}:\d{1,2}.*?$', re.MULTILINE) # 匹配所有含时间戳的行,去除每行首尾空白后返回 timestamp_lines = [line.strip() for line in timestamp_pattern.findall(description) if line.strip()] # 输出结果,也可以根据需求返回该列表 for line in timestamp_lines: print(line) print()
可选扩展:拆分时间戳和对应文本
如果需要把时间戳和描述文本单独拆分存储,可以用下面的逻辑:
def get_links(description): timestamp_lines = [] # 匹配行内任意位置的时间戳,同时捕获时间戳和前后文本 split_pattern = re.compile(r'(.*?)(\d{1,2}:\d{2}(?::\d{2})?)(.*)', re.MULTILINE) for match in split_pattern.finditer(description): pre_text, timestamp, post_text = match.groups() # 合并文本,去除多余空白 full_text = f"{pre_text.strip()} {post_text.strip()}".strip() timestamp_lines.append({ "timestamp": timestamp, "content": full_text }) # 打印结果示例 for item in timestamp_lines: print(f"时间戳:{item['timestamp']},描述:{item['content']}") print()
内容的提问来源于stack exchange,提问作者MarkWP
相关产品推荐
相关产品推荐

