如何用Python将特定对话文件转为列并保存为CSV?
大体积对话文本转CSV高效实现方案
实现思路
针对大文件场景,必须采用逐行流式处理,避免一次性加载全量数据到内存,从根源上控制内存占用,保证处理效率。
Python 代码实现
import csv def dialogue_to_csv(input_file, output_file): current_speaker = None current_content = [] # 双文件上下文管理器,同时处理输入输出 with open(input_file, 'r', encoding='utf-8') as in_f, \ open(output_file, 'w', encoding='utf-8', newline='') as out_f: csv_writer = csv.writer(out_f) # 写入CSV表头 csv_writer.writerow(['发言者', '对话内容']) for line in in_f: stripped_line = line.strip() # 跳过空行 if not stripped_line: continue # 识别发言者行(以>开头) if stripped_line.startswith('>'): # 若存在未写入的上一段对话,先写入CSV if current_speaker and current_content: full_content = '\n'.join(current_content) csv_writer.writerow([current_speaker, full_content]) current_content = [] # 提取发言者名称 current_speaker = stripped_line[1:] else: # 收集对话内容行 current_content.append(stripped_line) # 处理最后一段未写入的对话 if current_speaker and current_content: full_content = '\n'.join(current_content) csv_writer.writerow([current_speaker, full_content]) # 使用示例 # dialogue_to_csv('你的输入文件.txt', '输出结果.csv')
代码优势
- 低内存占用:逐行读取+即时写入,仅缓存当前对话片段,GB级文件也能稳定处理。
- 格式兼容:自动合并多行对话内容,CSV单元格内保留换行格式,符合阅读习惯。
- 性能高效:依赖Python内置模块,无额外依赖,处理速度接近文件IO极限。
转换后CSV内容示例
| 发言者 | 对话内容 |
|---|---|
| bernardo11_5 | 你这岗亭值守期间太平吗? |
| francisco11_5 | 连只老鼠动的动静都没有。 |
| bernardo11_6 | 行,晚安。 要是碰到霍雷肖和马塞勒斯——跟我轮值的那俩家伙,让他们快点过来。 |
| francisco11_6 | 我好像听见他们的声音了。——站住,谁在那儿? |
| horatio11_1 | 是这片地界的朋友。 |
| marcellus11_1 | 是丹麦王的臣民。 |
内容的提问来源于stack exchange,提问作者avi007
相关产品推荐
相关产品推荐

