You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python将特定对话文件转为列并保存为CSV?

大体积对话文本转CSV高效实现方案

实现思路

针对大文件场景,必须采用逐行流式处理,避免一次性加载全量数据到内存,从根源上控制内存占用,保证处理效率。

Python 代码实现

import csv

def dialogue_to_csv(input_file, output_file):
    current_speaker = None
    current_content = []
    
    # 双文件上下文管理器,同时处理输入输出
    with open(input_file, 'r', encoding='utf-8') as in_f, \
         open(output_file, 'w', encoding='utf-8', newline='') as out_f:
        
        csv_writer = csv.writer(out_f)
        # 写入CSV表头
        csv_writer.writerow(['发言者', '对话内容'])
        
        for line in in_f:
            stripped_line = line.strip()
            # 跳过空行
            if not stripped_line:
                continue
            # 识别发言者行(以>开头)
            if stripped_line.startswith('>'):
                # 若存在未写入的上一段对话,先写入CSV
                if current_speaker and current_content:
                    full_content = '\n'.join(current_content)
                    csv_writer.writerow([current_speaker, full_content])
                    current_content = []
                # 提取发言者名称
                current_speaker = stripped_line[1:]
            else:
                # 收集对话内容行
                current_content.append(stripped_line)
        
        # 处理最后一段未写入的对话
        if current_speaker and current_content:
            full_content = '\n'.join(current_content)
            csv_writer.writerow([current_speaker, full_content])

# 使用示例
# dialogue_to_csv('你的输入文件.txt', '输出结果.csv')

代码优势

  • 低内存占用:逐行读取+即时写入,仅缓存当前对话片段,GB级文件也能稳定处理。
  • 格式兼容:自动合并多行对话内容,CSV单元格内保留换行格式,符合阅读习惯。
  • 性能高效:依赖Python内置模块,无额外依赖,处理速度接近文件IO极限。

转换后CSV内容示例

发言者对话内容
bernardo11_5你这岗亭值守期间太平吗?
francisco11_5连只老鼠动的动静都没有。
bernardo11_6行,晚安。
要是碰到霍雷肖和马塞勒斯——跟我轮值的那俩家伙,让他们快点过来。
francisco11_6我好像听见他们的声音了。——站住,谁在那儿?
horatio11_1是这片地界的朋友。
marcellus11_1是丹麦王的臣民。

内容的提问来源于stack exchange,提问作者avi007

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 15:36:08