You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将多层嵌套JSON扁平化为单行列结构的Pandas DataFrame

将嵌套JSON的stats字段扁平化为DataFrame的同一行列

问题场景

现有如下嵌套JSON数据,需要将其扁平化为Pandas DataFrame并导出为CSV:

j = [
    {
        "id": 401281949,
        "teams": [
            {
                "school": "Louisiana Tech",
                "conference": "Conference USA",
                "homeAway": "away",
                "points": 34,
                "stats": [
                    {"category": "rushingTDs", "stat": "1"},
                    {"category": "puntReturnYards", "stat": "24"},
                    {"category": "puntReturnTDs", "stat": "0"},
                    {"category": "puntReturns", "stat": "3"},
                ],
            }
        ],
    }
]

执行pd.json_normalize(j, record_path=['teams'])后,stats仍为数组列;若通过explode再展开的方式,会生成多行数据,无法将每个stats字段作为独立列保留在同一行。

解决方案

可以通过将stats数组转换为字典后展开的方式,实现二次扁平化:

import pandas as pd

# 扁平化teams层,同时保留上层的id字段
df = pd.json_normalize(j, record_path=['teams'], meta=['id'])

# 将stats数组转换为{category: stat}格式的字典
df['stats'] = df['stats'].apply(lambda stats_list: {item['category']: item['stat'] for item in stats_list})

# 展开字典为独立列并合并到原DataFrame
df = df.join(pd.DataFrame(df.pop('stats').tolist()))

# 查看结果
print(df)

代码说明

  1. 扁平化teams层:利用pd.json_normalize的meta参数,在扁平化teams的同时保留上层的id字段,避免丢失关联数据。
  2. 转换stats为字典:将每个stats数组转换成键为统计类别、值为对应数据的字典,为后续展开列做准备。
  3. 展开字典为列:将字典列转换为DataFrame后,通过join合并到原数据中,实现所有统计项在同一行展示。

执行后得到的DataFrame结构如下:

school      conference homeAway points        id rushingTDs puntReturnYards puntReturnTDs puntReturns
0  Louisiana Tech  Conference USA      away     34  401281949          1              24             0           3

内容的提问来源于stack exchange,提问作者Josh Johnson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 12:05:21