You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何展平DataFrame的API_response列中嵌套JSON对象并提取指定字段?

解决嵌套JSON字符串列的展平与字段提取问题

你遇到的TypeError: string indices must be integers错误,核心原因是你的API_response列存储的是JSON格式的字符串,而非已经解析的Python字典/列表对象,pd.json_normalize()无法直接处理字符串类型的输入,所以才会触发报错。

下面是完整的解决方案,分步骤帮你提取所需字段并合并回原DataFrame:

步骤1:导入必要的库

import pandas as pd
import json

步骤2:解析JSON字符串为Python字典

首先需要把API_response列中的每个JSON字符串转换成Python可识别的字典对象,同时处理可能的空值避免报错:

# 解析JSON字符串,空值或空字符串转为空字典
df['API_response'] = df['API_response'].apply(
    lambda x: json.loads(x) if pd.notna(x) and x.strip() != '' else {}
)

步骤3:展平嵌套JSON并提取目标字段

使用pd.json_normalize()展平嵌套结构,指定record_path定位到resultater列表,同时保留顶层的kategori字段:

# 展平嵌套JSON,提取目标嵌套字段
flattened_df = pd.json_normalize(
    data=df['API_response'],
    record_path='resultater',  # 定位到嵌套的resultater列表
    meta=['kategori'],         # 保留顶层的kategori字段
    errors='ignore'            # 忽略解析过程中的错误
)

# 筛选出你需要的字段:kategori、adresse.id、adresse.adgangsadresseid
target_fields = flattened_df[['kategori', 'adresse.id', 'adresse.adgangsadresseid']]

步骤4:合并到原DataFrame

将提取出的字段合并回原DataFrame,同时可以选择移除原API_response列:

# 合并数据,丢弃原API_response列
final_df = pd.concat([df.drop('API_response', axis=1), target_fields], axis=1)

额外说明:处理多元素的resultater列表

如果你的部分行中resultater是包含多个对象的列表,上述代码会将每个列表元素展开为单独的行。如果你希望保留原行结构,只取列表中的第一个元素,可以在解析后先处理:

# 只保留resultater列表的第一个元素(如果列表存在)
df['API_response'] = df['API_response'].apply(
    lambda x: {**x, 'resultater': x.get('resultater', [])[0]} if x.get('resultater') else x
)

执行完上述代码后,你的原DataFrame就会新增kategori、adresse.id和adresse.adgangsadresseid三列,完全满足你的需求。

内容的提问来源于stack exchange,提问作者mantasbacys

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 10:17:43