嵌套JSON转扁平化DataFrame:自定义问题字段列映射实现求助
问题:嵌套JSON转扁平化DataFrame
现有嵌套格式JSON数据,包含分页信息及registrants数组:
- 每个
registrant对象包含基础信息(如姓名、邮箱等) - 同时包含
custom_questions数组,数组内每个元素有title(自定义问题名称)和value(对应回答)
需求是将custom_questions中的title作为DataFrame的列名,对应的value作为列值,实现扁平化结构。尝试以下代码未得到预期结果:
df = pd.json_normalize(json.loads(pat.explode("custom_questions").to_json(orient="records")))
解决方案
步骤1:提取核心数据
先从原始JSON中取出registrants数组(假设原始JSON已加载到变量data中):
import pandas as pd import json registrants = data['registrants']
步骤2:处理每个注册者数据
遍历每个registrant,将custom_questions转换为键值对字典,再与基础信息合并:
processed_data = [] for reg in registrants: # 提取基础信息(排除custom_questions字段) base_info = {key: val for key, val in reg.items() if key != 'custom_questions'} # 将custom_questions转为{title: value}的字典 custom_fields = {cq['title']: cq['value'] for cq in reg['custom_questions']} # 合并基础信息和自定义字段 processed_data.append({**base_info, **custom_fields})
步骤3:生成扁平化DataFrame
直接将处理后的列表转为DataFrame:
df = pd.DataFrame(processed_data)
原代码问题说明
explode("custom_questions")会将每个custom_questions元素拆分为单独行,导致同一个注册者对应多条记录,而非将自定义问题展开为列。上述方法通过字典合并,能确保每个注册者对应一行,自定义问题作为列直接展开。
内容的提问来源于stack exchange,提问作者Megan
相关产品推荐
相关产品推荐

