如何用Python去除转录文本中50词内重复出现的时间戳?
Python 实现时间戳过滤方案
核心思路
- 识别时间戳:用正则匹配文本中的时间戳(默认适配
[HH:MM:SS]或[MM:SS]格式,可按需调整) - 拆分文本块:把整个转录文本拆成「时间戳 + 后续内容」的片段集合
- 统计单词间隔:计算相邻时间戳对应内容的单词数,仅保留与上一个保留时间戳间隔≥50词的时间戳
- 重构文本:将筛选后的时间戳和对应内容重新拼接,输出处理结果
代码实现
import re def filter_timestamps(input_file, output_file, min_word_gap=50): # 读取原始转录文本 with open(input_file, 'r', encoding='utf-8') as f: text = f.read() # 匹配时间戳(可根据实际格式修改正则表达式) timestamp_pattern = re.compile(r'(\[\d{1,2}:\d{2}(:\d{2})?\])') # 拆分文本为时间戳和对应内容的片段 parts = timestamp_pattern.split(text) # 整理成[(时间戳, 对应内容), ...]的结构 timestamp_blocks = [] for i in range(1, len(parts), 3): ts = parts[i] content = parts[i+1].strip() if (i+1 < len(parts)) else '' timestamp_blocks.append((ts, content)) if not timestamp_blocks: print("未检测到时间戳格式内容") return # 筛选符合间隔要求的时间戳 filtered_blocks = [timestamp_blocks[0]] # 累计上一个保留时间戳后的单词数 total_words = len(filtered_blocks[0][1].split()) for ts, content in timestamp_blocks[1:]: current_words = len(content.split()) # 检查累计单词数是否达到间隔要求 if total_words >= min_word_gap: filtered_blocks.append((ts, content)) total_words = current_words else: total_words += current_words # 重构处理后的文本 result = [] for ts, content in filtered_blocks: result.append(f"{ts} {content}") final_text = '\n'.join(result) # 写入输出文件 with open(output_file, 'w', encoding='utf-8') as f: f.write(final_text) print(f"处理完成,结果已保存至{output_file}") # 使用示例 if __name__ == "__main__": filter_timestamps("原始转录文本.txt", "过滤后文本.txt")
关键细节说明
- 时间戳格式适配:如果你的时间戳是其他样式(比如
00:00:00不带方括号),直接修改timestamp_pattern的正则表达式即可 - 单词计数优化:当前按空格分割统计单词数,若需要排除标点干扰,可在拆分前添加清洗逻辑:
content = re.sub(r'[^\w\s]', '', content) - 边界场景处理:若第一个时间戳后的内容不足50词,后续时间戳的单词数会持续累加,直到总间隔达标才保留下一个时间戳
内容的提问来源于stack exchange,提问作者AC123
相关产品推荐
相关产品推荐

