行首匹配日期条件下多行文本换行替换方法咨询
多行聊天记录拼接处理方案
问题背景
现有格式为聊天导出记录的文本文件,绝大多数行首带有可正则匹配的日期时间格式,要求行首未匹配到日期格式的行,删除该行前的换行符,将内容拼接至前一行末尾。
原始文件样例:
3/11/21, 6:13 PM - Gil: X1000 3/11/21, 6:15 PM - Sergio: <Media omitted> 3/11/21, 6:19 PM - Sergio: X400 3/11/21, 6:20 PM - Sergio: Los amigos de vonzo en Francia: 1. La Tóxica 2. El brujo vodoo 3. El/La Zoofilic@ 3/11/21, 6:20 PM - Sergio: :V 3/11/21, 6:21 PM - Joan :V: JAJAJAJAJA
预期输出效果:
3/11/21, 6:13 PM - Gil: X1000 3/11/21, 6:15 PM - Sergio: <Media omitted> 3/11/21, 6:19 PM - Sergio: X400 3/11/21, 6:20 PM - Sergio: Los amigos de vonzo en Francia: 1. La Tóxica 2. El brujo vodoo 3. El/La Zoofilic@ 3/11/21, 6:20 PM - Sergio: :V 3/11/21, 6:21 PM - Joan :V: JAJAJAJAJA
实现思路
逐行读取时不要直接输出,新增一个变量缓存上一行的内容:
- 读取当前行后先判断行首是否匹配日期格式
- 匹配则输出缓存的上一行,再把当前行存入缓存
- 不匹配则把当前行去掉空白后拼接到缓存的上一行末尾
- 所有行读取完成后输出最后缓存的内容即可
完整实现代码
import re # 匹配行首日期的正则,适配你提供的日期格式 date_pattern = re.compile(r'^\d{1,2}/\d{1,2}/\d{2,4},\s\d{1,2}:\d{2}\s[AP]M') # 缓存上一行内容 prev_line = '' with open(self.fileName, encoding="utf8", errors='replace') as input_file, open('输出文件路径.txt', 'w', encoding='utf8') as output: for line in input_file: # 去除当前行的换行符和首尾空白 stripped_line = line.strip() # 跳过空行 if not stripped_line: continue if date_pattern.match(stripped_line): # 当前行是新的聊天行,先输出上一行缓存 if prev_line: output.write(prev_line + '\n') prev_line = stripped_line else: # 当前行是上一条聊天的换行内容,直接拼接 prev_line += ' ' + stripped_line # 输出最后一条缓存的内容 if prev_line: output.write(prev_line + '\n')
调整说明
如果不需要拼接时插入空格分隔,把prev_line += ' ' + stripped_line中的空格删除即可。
内容的提问来源于stack exchange,提问作者Ivan
相关产品推荐
相关产品推荐

