You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用pd.json_normalize解析含嵌套列表的JSON文件?

解决嵌套JSON列表的表格解析问题

针对JSON中attributes和tags这类嵌套列表无法被pd.json_normalize自动解析的问题,以下是两种实用的处理方案:

方法1:将嵌套列表展开为多行,保留原字段关联

如果需要把每个列表项单独作为一行,同时保留原数据的uuid、timestamp等标识字段,可以用explode配合字典转列的方式:

import pandas as pd
import json

# 加载JSON数据
with open('your_data.json', 'r') as f:
    data = json.load(f)

# 先标准化顶层字段,生成基础DataFrame
df_base = pd.json_normalize(data)

# 处理attributes字段:展开列表为多行,再拆分字典为列
df_attributes = df_base.explode('attributes').reset_index(drop=True)
df_attributes = pd.concat([
    df_attributes.drop('attributes', axis=1),
    df_attributes['attributes'].apply(pd.Series)
], axis=1)

# 同理处理tags字段(按需选择是否同时展开,注意多列表展开会产生笛卡尔积)
df_tags = df_base.explode('tags').reset_index(drop=True)
df_tags = pd.concat([
    df_tags.drop('tags', axis=1),
    df_tags['tags'].apply(pd.Series)
], axis=1)

方法2:将嵌套列表的键值对转换为表格列

如果希望把嵌套列表中的键值对直接转为表格的列(比如attributes里的key作为列名,value作为对应值),可以先预处理数据结构:

import pandas as pd
import json

with open('your_data.json', 'r') as f:
    data = json.load(f)

# 自定义函数:将attributes列表转为字典并合并到原数据项
def flatten_attributes(item):
    attr_dict = {attr['key']: attr['value'] for attr in item['attributes']}
    return {**item, **attr_dict}

# 自定义函数:处理tags列表(根据实际结构调整键值映射逻辑)
def flatten_tags(item):
    tag_dict = {tag['name']: tag['category'] for tag in item['tags']}
    return {**item, **tag_dict}

# 依次处理嵌套字段
processed_data = [flatten_attributes(item) for item in data]
processed_data = [flatten_tags(item) for item in processed_data]

# 标准化处理后的数据,并移除原嵌套字段
df = pd.json_normalize(processed_data)
df = df.drop(['attributes', 'tags'], axis=1)

注意事项

  • 如果嵌套列表中存在重复的键(比如同一个attributes里有两个age),后出现的键值会覆盖前者,需提前做去重或聚合处理
  • 同时展开多个嵌套列表时,explode会生成笛卡尔积,需确认业务逻辑是否允许
  • 请根据你的JSON实际字段名调整代码中的键名(比如key/value或name/category)

内容的提问来源于stack exchange,提问作者Nicolas Fabre

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 00:22:36