You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何读取嵌套贷款JSON文件并转换为Pandas DataFrame用于分析

处理嵌套JSON贷款数据并转换为结构化DataFrame

核心思路

针对嵌套JSON格式的贷款数据,利用Pandas的json_normalize工具扁平化嵌套结构,精准提取目标字段(loanId、TransactionStatus、AccountType),生成适合拒贷推理分析的结构化DataFrame。


步骤1:示例嵌套JSON数据

假设你的贷款数据文件loan_data.json结构如下(包含多层嵌套):

[
  {
    "loanId": "LN12345",
    "customerInfo": {
      "accountDetails": {
        "AccountType": "Savings",
        "accountNumber": "ACC789"
      }
    },
    "transaction": {
      "TransactionStatus": "Approved",
      "amount": 5000
    }
  },
  {
    "loanId": "LN67890",
    "customerInfo": {
      "accountDetails": {
        "AccountType": "Current",
        "accountNumber": "ACC101"
      }
    },
    "transaction": {
      "TransactionStatus": "Rejected",
      "amount": 10000
    }
  }
]

步骤2:读取并扁平化数据

使用Python的json模块读取文件,再通过pd.json_normalize处理嵌套结构:

import pandas as pd
import json

# 读取JSON文件
with open('loan_data.json', 'r') as f:
    loan_data = json.load(f)

# 扁平化嵌套结构并提取目标字段
df = pd.json_normalize(
    loan_data,
    meta=[
        'loanId',
        ['transaction', 'TransactionStatus'],
        ['customerInfo', 'accountDetails', 'AccountType']
    ]
)

# 重命名列名,简化字段名
df.columns = ['loanId', 'TransactionStatus', 'AccountType']

# 查看结构化结果
print(df)

输出结果:

loanId TransactionStatus AccountType
0  LN12345           Approved      Savings
1  LN67890           Rejected      Current

步骤3:处理含数组的复杂嵌套

如果贷款数据包含多交易记录的数组嵌套(比如单贷款对应多笔交易),调整record_path参数指定数组字段:
示例JSON结构:

[
  {
    "loanId": "LN12345",
    "customerInfo": {
      "accountDetails": {
        "AccountType": "Savings"
      }
    },
    "transactions": [
      {"TransactionStatus": "Approved", "amount": 5000},
      {"TransactionStatus": "Pending", "amount": 2000}
    ]
  }
]

对应处理代码:

df = pd.json_normalize(
    loan_data,
    record_path='transactions',  # 以交易数组作为行数据源
    meta=[
        'loanId',
        ['customerInfo', 'accountDetails', 'AccountType']
    ]
)

# 筛选并整理目标字段
final_df = df[['loanId', 'TransactionStatus', 'AccountType']]
print(final_df)

输出结果:

loanId TransactionStatus AccountType
0  LN12345           Approved      Savings
1  LN12345            Pending      Savings

实用注意事项

  • 全量扁平化后筛选:如果不确定嵌套层级,可直接执行pd.json_normalize(loan_data)生成全量扁平化DataFrame,再通过列名筛选目标字段(如df.filter(regex='loanId|TransactionStatus|AccountType'))。
  • 缺失值处理:嵌套字段可能存在缺失,可通过df.fillna('Unknown')填充默认值,或df.dropna(subset=['loanId'])删除关键字段缺失的行。
  • JSON格式校验:确保JSON文件无语法错误,否则会导致加载失败,可使用本地JSON校验工具提前验证。

内容的提问来源于stack exchange,提问作者SHESADEV SHA

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 11:17:04