You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python快速提取海量聊天机器人数据中的客户对话

高效提取聊天记录中客户对话与时间信息的Python实现

核心思路

针对数百万条聊天记录的高效处理,核心在于预编译正则表达式(避免重复编译开销)、逐行流式处理(减少内存占用),以及尽量使用底层高效的字符串操作。

实现步骤与代码

1. 预编译正则表达式

提前编译匹配聊天开始时间和客户消息的正则模式,这是提升批量处理效率的关键:

import re
from datetime import datetime

# 预编译正则模式
time_pattern = re.compile(r'Chat Started: (\w+ \w+ \d+ \d+ \d+:\d+:\d+)')
customer_msg_pattern = re.compile(r'(Customer\d+): (.+)')

2. 单条聊天记录处理函数

编写函数处理单条聊天会话文本,提取时间和客户对话:

def process_chat_session(chat_text):
    # 提取聊天开始时间并格式化
    time_match = time_pattern.search(chat_text)
    if not time_match:
        return None
    raw_time = time_match.group(1)
    formatted_time = datetime.strptime(raw_time, '%A %B %d %Y %H:%M:%S').strftime('%d-%b-%Y %H:%M:%S')
    
    # 提取所有客户消息
    customer_msgs = {}
    for match in customer_msg_pattern.finditer(chat_text):
        customer_id = match.group(1)
        msg = match.group(2).strip()
        if customer_id not in customer_msgs:
            customer_msgs[customer_id] = set()
        customer_msgs[customer_id].add(msg)
    
    # 按要求格式输出结果
    results = []
    for cust_id, msgs in customer_msgs.items():
        results.append(f"{cust_id}: {msgs} Time: {'{' + formatted_time + '}'}")
    return results

3. 批量处理优化

对于数百万条记录,建议逐行读取文件(而非一次性加载全部内容),并按聊天会话分割处理:

def batch_process_chats(file_path):
    with open(file_path, 'r', encoding='utf-8') as f:
        current_session = []
        for line in f:
            line = line.strip()
            if not line:
                continue
            # 检测新会话开始(假设每条会话以"Chat Started:"开头)
            if line.startswith('Chat Started:'):
                if current_session:
                    # 处理上一个会话
                    session_text = ' '.join(current_session)
                    results = process_chat_session(session_text)
                    if results:
                        for res in results:
                            print(res)  # 可替换为写入输出文件等操作
                current_session = [line]
            else:
                current_session.append(line)
        # 处理最后一个会话
        if current_session:
            session_text = ' '.join(current_session)
            results = process_chat_session(session_text)
            if results:
                for res in results:
                    print(res)

性能优化建议

  • 多进程处理:如果聊天记录分散在多个文件中,可使用multiprocessing模块并行处理不同文件,充分利用CPU多核。
  • 减少内存开销:避免将所有聊天内容存入内存,始终采用流式逐行处理。
  • 正则优化:确保正则表达式尽可能简洁(比如避免不必要的捕获组),进一步提升匹配速度。

内容的提问来源于stack exchange,提问作者Chandra Sekhar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 23:45:42