You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对嵌套JSON列表执行Normalize/Flatten?选JSON flatten还是pandas?

嵌套JSON列表扁平化处理方案

需求要点

你手头的嵌套JSON列表需要做这些处理:

  • 顶层的index、type、id、score直接保留
  • source里的嵌套字段(比如user.resource_uuid、resource.resource_uuid)要拉到同一层级
  • in_context下的所有字段全部拉平
  • response字段是JSON格式的字符串,得解析后展开成带序号的字段(比如recipientid1、text1)

方案选择:Pandas Normalize + 自定义处理

光用json-flatten库只能处理常规的嵌套JSON对象,但你这需求特殊在response是JSON字符串而非嵌套对象,还得额外解析展开,所以用Pandas结合自定义函数更顺手,能一次性搞定所有操作。

具体代码实现

import pandas as pd
import json

# 你的原始数据列表
data = [
    {'index': 'exp-000005',
     'type': '_doc',
     'id': 'jdaksjdlkj',
     'score': 9.502488,
     'source': {'user': {'resource_uuid': '123'},
                'verb': 'REPLIED',
                'resource': {'resource_uuid': '/home'},
                'timestamp': '2022-01-20T08:14:00+00:00',
                'with_result': {},
                'in_context': {'screen_width': '3440',
                               'screen_height': '1440',
                               'build_version': '7235',
                               'clientid': '123',
                               'question': 'Hallo',
                               'request_time': '403',
                               'status': 'success',
                               'response': '[]',
                               'language': 'de'}}},
    {'index': 'exp-000005',
     'type': '_doc',
     'id': 'dddddd',
     'score': 9.502488,
     'source': {'user': {'resource_uuid': '44444'},
                'verb': 'REPLIED',
                'resource': {'resource_uuid': '/home'},
                'timestamp': '2022-01-20T08:14:10+00:00',
                'with_result': {},
                'in_context': {'screen_width': '3440',
                               'screen_height': '1440',
                               'build_version': '7235',
                               'clientid': '345',
                               'question': 'Ich brauche Hilfe',
                               'request_time': '111',
                               'status': 'success',
                               'response': '[{"recipientid":"789", "text":"Bitte sehr."}, {"recipientid":"888", "text":"Kann ich Ihnen noch mit etwas anderem behilflich sein?"}]',
                               'language': 'de'}}},
    {'index': 'exp-000005',
     'type': '_doc',
     'id': 'jdhdgs',
     'score': 9.502488,
     'source': {'user': {'resource_uuid': '333'},
                'verb': 'REPLIED',
                'resource': {'resource_uuid': '/home'},
                'timestamp': '2022-01-20T08:14:19+00:00',
                'with_result': {},
                'in_context': {'screen_width': '3440',
                               'screen_height': '1440',
                               'build_version': '7235',
                               'clientid': '007',
                               'question': 'Zertifikate',
                               'request_time': '121',
                               'status': 'success',
                               'response': '[{"recipientid":"345", "text":"Künstliche Intelligenz"}, {"recipientid":"123", "text":"Kann ich Ihnen noch mit etwas anderem behilflich sein?"}]',
                               'language': 'de'}}}
]

# 第一步:用json_normalize拉平顶层嵌套
df = pd.json_normalize(data, sep='_')

# 第二步:解析response里的JSON字符串并展开
def parse_response(row):
    try:
        # 把字符串转成列表
        resp_list = json.loads(row['source_in_context_response'])
        # 给每个元素加序号,生成新字段
        for idx, item in enumerate(resp_list, 1):
            row[f'recipientid{idx}'] = item.get('recipientid')
            row[f'text{idx}'] = item.get('text')
        # 去掉原response字段(不需要可以保留)
        row.pop('source_in_context_response')
    except json.JSONDecodeError:
        # 解析失败就跳过
        pass
    return row

# 对每一行应用处理函数
df = df.apply(parse_response, axis=1)

# 第三步:重命名字段,匹配你的需求格式
# 注意:JSON不允许同一层级有重复key,所以把两个resource_uuid区分开了
df = df.rename(columns={
    'source_user_resource_uuid': 'user_resource_uuid',
    'source_resource_resource_uuid': 'resource_resource_uuid',
    'source_verb': 'verb',
    'source_timestamp': 'timestamp',
    'source_with_result': 'with_result',
    'source_in_context_screen_width': 'screen_width',
    'source_in_context_screen_height': 'screen_height',
    'source_in_context_build_version': 'build_version',
    'source_in_context_clientid': 'clientid',
    'source_in_context_question': 'question',
    'source_in_context_request_time': 'request_time',
    'source_in_context_status': 'status',
    'source_in_context_language': 'language'
})

# 转成JSON列表输出
flattened_data = df.to_dict('records')
print(flattened_data)

额外说明

  1. 关于重复key:你期望的结果里有两个resource_uuid,但JSON规范不允许同一对象存在重复key,最终只会保留最后一个值,所以代码里把它们重命名成了user_resource_uuid和resource_resource_uuid,如果业务必须保留同名,得先确认逻辑,但不建议这么做。
  2. json-flatten的不足:如果用json-flatten库,它没法自动解析response里的JSON字符串,还得额外写代码处理,不如Pandas一步到位。
  3. 扩展性:要是后续嵌套结构变了,Pandas的json_normalize可以通过record_path和meta参数灵活调整,适配更复杂的场景。

内容的提问来源于stack exchange,提问作者threxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 13:14:49