You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用pd.json_normalize将嵌套JSON扁平化适配CSV存储?

用pd.json_normalize处理API数据并拆分tierList字段

完全可以用pd.json_normalize实现你的需求,先解决你遇到的核心问题:

为什么record_path='memberRiskData'会报错?

pd.json_normalize的record_path参数必须指向数组类型的字段(可迭代的列表/数组),如果你的API返回结构里memberRiskData是memberList数组中的单个对象(而非数组),直接指定它作为record_path就会报错——因为工具找不到可展开的迭代对象,这就是你只能用record_path='memberList'的原因。

拆分tierList为独立列的实现步骤

假设你的API返回数据结构类似这样(根据描述推导):

{
  "bin": "654321",
  "api_version": "v1",
  "memberList": [
    {
      "memberId": "M001",
      "memberRiskData": {
        "tierList": {
          "TierA": 0.95,
          "TierB": 0.7,
          "TierC": 0.3
        },
        "riskScore": 85
      }
    }
  ]
}

步骤1:展开memberList并保留顶层字段

先通过record_path='memberList'展开数组,同时用meta带上bin等需要保留的顶层字段:

import pandas as pd

# 替换为你的API返回数据
api_response = {"bin": "654321", ...}

# 展开memberList,保留顶层的bin字段
df = pd.json_normalize(
    api_response,
    record_path="memberList",
    meta=["bin"]  # 可添加其他顶层字段,比如"api_version"
)

步骤2:拆分tierList字典为独立列

此时df里会有memberRiskData.tierList列(值为字典),直接用pd.json_normalize拆分:

# 拆分tierList列
tier_columns = pd.json_normalize(df["memberRiskData.tierList"])

# 合并拆分后的列到原DataFrame,删除原tierList列
final_df = pd.concat([df.drop("memberRiskData.tierList", axis=1), tier_columns], axis=1)

步骤3:整理为以bin为一行的结构

如果一个bin对应多条member数据,可通过聚合实现单bin单行:

final_df = final_df.groupby("bin").agg("first").reset_index()

工具推荐

  • 首选pandas:pd.json_normalize+concat的组合完全覆盖需求,上手成本低,适配Python工作流。
  • jq(命令行工具):适合在Python脚本外预处理JSON,用.[].memberList[].memberRiskData.tierList语法可快速提取并展开字段。
  • Dask:如果数据量极大导致pandas内存不足,Dask的json_normalize支持分布式处理,用法与pandas几乎一致。

内容的提问来源于stack exchange,提问作者user1982778

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 01:27:09