You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则从带时间戳的聊天记录字符串中提取对话内容

解决方案

推荐使用分步处理的方案,相比复杂正则可读性更高、容错性更强,具体实现如下:

分步处理实现(推荐)

import re

string = """(2021-07-02 01:00:00 AM BST)  
---  
syl.hs has joined the conversation  
  
  

(2021-07-02 01:00:23 AM BST)  
---  
e.wang  
Good Morning
How're you?
  
  
  

(2021-07-02 01:05:11 AM BST)  
---  
wk.wang  
Hi, I'm Good.  
  
  

(2021-07-02 01:08:01 AM BST)  
---  
perter.derrek   
we got the update on work. 
It will get complete by next week.

(2021-07-15 08:59:41 PM BST)  
---  
ad.ft has left the conversation  
  
  
  
  

---  

* * *"""

comments = []
# 按时间戳+分隔线拆分每个聊天块
blocks = re.split(r'\(\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2} [AP]M BST\)\s*---\s*', string)
for block in blocks:
    block = block.strip()
    if not block:
        continue
    # 过滤入群、退群通知
    if 'has joined' in block or 'has left' in block:
        continue
    # 按行拆分,过滤空行
    lines = [line.strip() for line in block.splitlines() if line.strip()]
    if len(lines) < 2:
        continue
    # 合并除用户名外的所有内容行,去除多余空格
    content = ' '.join(lines[1:])
    comments.append(content)

print(comments)

运行后输出与预期完全一致:

["Good Morning How're you?", "Hi, I'm Good.", 'we got the update on work. It will get complete by next week.']


原正则问题与修正方案

你原来的正则存在以下问题:

  • 时区硬编码为GMT,但实际日志时区为BST,导致匹配失效
  • 未过滤入群、退群的通知类消息
  • 未区分用户名与后续聊天内容,会把用户名也提取到结果中
  • 没有处理多行内容的换行与多余空白

如果坚持使用单条正则匹配,可以用调整后的写法:

import re
pattern = r'---\s*\n(?!.*(?:joined|left))[^\n]+\n(.*?)(?=\n\(\d{4}-\d{2}-\d{2}|\Z)'
matches = re.findall(pattern, string, re.DOTALL)
comments = [' '.join([line.strip() for line in content.splitlines() if line.strip()]) for content in matches]

内容的提问来源于stack exchange,提问作者Learner

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 18:45:03