You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何以Pythonic方式遍历嵌套列表,实现JSON Payload归一化

处理嵌套列表的JSON归一化问题

原始Payload

payload = [{
  "identification": {
    "corporationName": "Test",
    "code": 1234
  },
  "locations": [
     {
      "postalAddress": {
        "address": "St Test",
        "country": "US"
      },
      "availability" : {
          "standards": [
            {
              "weekday": "SUNDAY",
              "openingTime": "string",
              "closingTime": "string"
            }
          ],
          "isPublicAccessAllowed": false
      },
      "phones": [
        {
          "countryCallingCode": "55",
          "areaCode": "19",
          "number": "35721199"
        }
      ]
    }
  ]
}]

现有代码

def generate_values(r: dict):
        return (
            ('companyname', get(r, 'identification.corporationName', valueType=str)),
            ('companycode', get(r, 'identification.code', valueType=str)),
        )
    
_ = [dict_test[k].append(v) for r in payload for k, v in generate_values(r)]

注:get为自定义的JSON值提取方法

需求

需要遍历Payload中的嵌套列表(如locations),同时提取父级及嵌套层级的字段,生成包含多层级键值对的集合,预期的键值对示例如下:

return (
    ('companyname', get(r, 'identification.corporationName', valueType=str)),
    ('companycode', get(r, 'identification.code', valueType=str)),
    ('addressstreet', get(nested_list_payload, 'postalAddress.address',valueType=str)),
    ('adresscountry', get(nested_list_payload, 'postalAddress.country',valueType=str)),
    ('anothervalue', get(nested_nested_list, ...))
)

解决方案

1. 修改generate_values函数,支持多层级数据传入

调整函数参数,同时接收父级对象和嵌套列表的子对象,在函数内部处理更深层次的嵌套列表(如standards、phones):

def generate_values(r: dict, loc: dict):
    # 提取父级(公司)字段
    yield ('companyname', get(r, 'identification.corporationName', valueType=str))
    yield ('companycode', get(r, 'identification.code', valueType=str))
    
    # 提取location层级字段
    yield ('addressstreet', get(loc, 'postalAddress.address', valueType=str))
    yield ('addresscountry', get(loc, 'postalAddress.country', valueType=str))
    yield ('public_access_allowed', get(loc, 'availability.isPublicAccessAllowed', valueType=bool))
    
    # 处理availability下的standards嵌套列表
    for standard in get(loc, 'availability.standards', default=[], valueType=list):
        yield ('weekday', get(standard, 'weekday', valueType=str))
        yield ('opening_time', get(standard, 'openingTime', valueType=str))
        yield ('closing_time', get(standard, 'closingTime', valueType=str))
    
    # 处理phones嵌套列表
    for phone in get(loc, 'phones', default=[], valueType=list):
        yield ('phone_country_code', get(phone, 'countryCallingCode', valueType=str))
        yield ('phone_area_code', get(phone, 'areaCode', valueType=str))
        yield ('phone_number', get(phone, 'number', valueType=str))

2. 调整遍历逻辑,覆盖嵌套列表

使用嵌套遍历,先遍历顶层的公司对象,再遍历每个公司下的locations列表项,将父级对象和子对象同时传入generate_values:

from collections import defaultdict

# 初始化字典,避免键不存在的KeyError
dict_test = defaultdict(list)

# 嵌套遍历顶层对象和locations列表
for r in payload:
    for loc in get(r, 'locations', default=[], valueType=list):
        for k, v in generate_values(r, loc):
            dict_test[k].append(v)

如果坚持使用列表推导式(可读性稍差),可以写成:

from collections import defaultdict
dict_test = defaultdict(list)

_ = [dict_test[k].append(v) for r in payload for loc in get(r, 'locations', default=[]) for k, v in generate_values(r, loc)]

说明

  • 遍历逻辑采用顶层对象 → locations子对象 → 更深层嵌套列表的层级,确保所有嵌套数据都被覆盖
  • generate_values使用yield生成器,逐一生成键值对,避免一次性生成大量元组占用内存
  • 对可能为空的嵌套列表(如locations、standards)添加default=[],防止遍历出错

内容的提问来源于stack exchange,提问作者CooperativistKid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 08:06:14