You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python Pandas将大型JSON文件转换为CSV?

用Python Pandas将嵌套JSON转换为CSV的解决方案

嘿,我来帮你搞定这个JSON转CSV的需求!你的JSON数据是嵌套数组结构:外层数组中的每个元素都是一段对话(包含多条用户消息),每条消息是带user/User、text、sent字段的字典。用Pandas处理的话,核心是先把嵌套的数据扁平化,再转成CSV,步骤如下:

完整代码示例

import json
import pandas as pd

# 1. 读取并加载JSON文件
with open('/Users/dsg281/Downloads/EmotionLines/Friends/friends_dev.json', 'r', encoding='utf-8') as f:
    data = json.load(f)

# 2. 扁平化嵌套数据:提取所有对话条目为一维列表
# 同时统一字段名(比如把"User"转成小写"user",避免列分裂)
flat_data = []
for conversation in data:
    for message in conversation:
        # 统一键名到小写,避免大小写导致的列不一致
        normalized_msg = {key.lower(): value for key, value in message.items()}
        flat_data.append(normalized_msg)

# 3. 转换为Pandas DataFrame
df = pd.DataFrame(flat_data)

# 4. 保存为CSV文件
df.to_csv('friends_dev_converted.csv', index=False, encoding='utf-8')

关键步骤解释

  • 读取JSON:用json.load()直接加载文件内容为Python嵌套列表,比Pandas的read_json更适配这种非标准的嵌套对话结构。
  • 扁平化数据:原始数据是「对话组→多条消息」的层级,我们需要把每一条独立消息作为CSV的一行,所以用双层循环把所有消息提取到一维列表中。
  • 统一字段名:你的示例数据里同时存在User(大写U)和user(小写u),这会导致Pandas生成两列,所以用字典推导式把所有键转成小写,保证字段一致性。
  • 保存CSV:index=False避免把DataFrame的索引写入CSV,encoding='utf-8'确保对话中的特殊字符正常保存。

测试示例数据

如果用你提供的示例JSON运行代码,最终的CSV会是这样的结构化内容:

usertextsent
PhoebeOh my God, hes lost it. Hes totally lost it.non-neutral
MonicaWhat?surprise
JoeyHey Estelle, listenneutral
EstelleWell! Well! Well! Joey Tribbiani! So you came back huh? Theysurprise

这样就完美把嵌套的对话数据转换成了易读的CSV格式啦!

内容的提问来源于stack exchange,提问作者Kum_R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:12:40