You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何清洗Cornell Movie-Dialogs Corpus数据集并转为CSV文件?

处理Cornell Movie-Dialogs Corpus文本并转换为CSV

我之前也折腾过这个数据集,这种特殊分隔符确实有点绕,不过拆解起来其实很简单,咱们一步步来解决:

1. 先理清数据结构

每条记录用+++$+++作为分隔符拆分后,字段依次是:

  • 行ID(比如L1045)
  • 用户ID(u0)
  • 电影ID(m0)
  • 角色名(BIANCA)
  • 对话内容(这正是你需要提取的核心部分)

2. 单条对话提取并转CSV

如果只是要把每条对话内容单独导出成CSV,用Python写几行代码就能搞定,直接用这个示例:

import csv

# 替换成你本地的原始文本文件路径
input_file = "movie_lines.txt"
# 输出的CSV文件路径
output_file = "cleaned_dialogues.csv"

with open(input_file, 'r', encoding='iso-8859-1') as infile, open(output_file, 'w', newline='', encoding='utf-8') as outfile:
    writer = csv.writer(outfile)
    # 写入CSV表头
    writer.writerow(["Dialogue"])
    
    for line in infile:
        line = line.strip()
        if not line:
            continue
        # 用分隔符拆分字段,取最后一个就是对话内容
        parts = line.split('+++$+++')
        dialogue = parts[-1].strip()
        # 写入CSV
        writer.writerow([dialogue])

3. 如果需要构建对话对(适合聊天机器人训练)

要是你训练聊天机器人需要“上下文-回复”的配对数据,这个数据集里的movie_conversations.txt记录了每段对话的行ID顺序,结合它来构建配对更合理:

import csv
import re

# 先把所有对话用行ID做索引存起来
line_id_to_dialogue = {}
with open("movie_lines.txt", 'r', encoding='iso-8859-1') as lines_file:
    for line in lines_file:
        line = line.strip()
        if not line:
            continue
        parts = line.split('+++$+++')
        line_id = parts[0].strip()
        dialogue = parts[-1].strip()
        line_id_to_dialogue[line_id] = dialogue

# 读取对话序列,生成上下文-回复配对
with open("movie_conversations.txt", 'r', encoding='iso-8859-1') as conv_file, open("dialogue_pairs.csv", 'w', newline='', encoding='utf-8') as outfile:
    writer = csv.writer(outfile)
    writer.writerow(["Context", "Response"])
    
    for line in conv_file:
        line = line.strip()
        if not line:
            continue
        # 提取括号里的对话ID列表
        match = re.search(r'\[(.*?)\]', line)
        if not match:
            continue
        line_ids_str = match.group(1)
        line_ids = [id.strip().strip("'") for id in line_ids_str.split(',')]
        
        # 把相邻的对话配对:前一句是上下文,后一句是回复
        for i in range(len(line_ids) - 1):
            context_id = line_ids[i]
            response_id = line_ids[i+1]
            if context_id in line_id_to_dialogue and response_id in line_id_to_dialogue:
                writer.writerow([line_id_to_dialogue[context_id], line_id_to_dialogue[response_id]])

小提醒

这个数据集的编码是iso-8859-1,读取的时候一定要指定这个编码,不然容易出现乱码问题。

内容的提问来源于stack exchange,提问作者LostAtlas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 07:07:45