如何用Python正则从带时间戳的聊天记录字符串中提取对话内容
解决方案
推荐使用分步处理的方案,相比复杂正则可读性更高、容错性更强,具体实现如下:
分步处理实现(推荐)
import re string = """(2021-07-02 01:00:00 AM BST) --- syl.hs has joined the conversation (2021-07-02 01:00:23 AM BST) --- e.wang Good Morning How're you? (2021-07-02 01:05:11 AM BST) --- wk.wang Hi, I'm Good. (2021-07-02 01:08:01 AM BST) --- perter.derrek we got the update on work. It will get complete by next week. (2021-07-15 08:59:41 PM BST) --- ad.ft has left the conversation --- * * *""" comments = [] # 按时间戳+分隔线拆分每个聊天块 blocks = re.split(r'\(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2} [AP]M BST\)\s*---\s*', string) for block in blocks: block = block.strip() if not block: continue # 过滤入群、退群通知 if 'has joined' in block or 'has left' in block: continue # 按行拆分,过滤空行 lines = [line.strip() for line in block.splitlines() if line.strip()] if len(lines) < 2: continue # 合并除用户名外的所有内容行,去除多余空格 content = ' '.join(lines[1:]) comments.append(content) print(comments)
运行后输出与预期完全一致:
["Good Morning How're you?", "Hi, I'm Good.", 'we got the update on work. It will get complete by next week.']
原正则问题与修正方案
你原来的正则存在以下问题:
- 时区硬编码为
GMT,但实际日志时区为BST,导致匹配失效 - 未过滤入群、退群的通知类消息
- 未区分用户名与后续聊天内容,会把用户名也提取到结果中
- 没有处理多行内容的换行与多余空白
如果坚持使用单条正则匹配,可以用调整后的写法:
import re pattern = r'---\s*\n(?!.*(?:joined|left))[^\n]+\n(.*?)(?=\n\(\d{4}-\d{2}-\d{2}|\Z)' matches = re.findall(pattern, string, re.DOTALL) comments = [' '.join([line.strip() for line in content.splitlines() if line.strip()]) for content in matches]
内容的提问来源于stack exchange,提问作者Learner
相关产品推荐
相关产品推荐

