You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

用Pythonic的DRY方式从嵌套列表与字典构建DataFrame

如何用更Pythonic且符合DRY原则的方式扁平化嵌套字典列表为DataFrame?

我有一个包含字典的列表,部分字典中还嵌套了字典列表,具体数据如下:

samples = [
    {"person": "A", "employed": True, "location": "East", 
    "donations": []},
    {"person": "B", "employed": False, "location": "West", 
    "donations": [
                 {"type": "cash", "date": "july 25, 2022", "value": 50},
                 {"type": "stocks", "date": "May 1, 2022", "value": 2500}
                 ]},
    {"person": "C", "employed": False, "location": "North",
    "donations": []}
]

我需要将这些数据处理为扁平化嵌套字典的DataFrame,使其具有指定的列头。我现有的代码存在大量重复,请问如何采用更Pythonic且符合DRY原则的实现方式?

现有代码如下:

data_list = [] ## initial empty list
for sample in samples:
    ## get data that is not nested
    person = sample.get("person")
    employed = sample.get("employed")
    location = sample.get("location")
    ## if there is not nested data, give specific values
    if len(sample.get("donations")) < 1:
        donation_type = donation_date =  donation_value = "NONE LISTED"
        ## build list
        data_list.append({"person": person,
                  "employed": employed,
                  "location": location,
                  "donation_type": donation_type,
                  "donation_date": donation_date,
                  "donation_value": donation_value,
                 })
    else:
        ## build same list but with provided values
        donations = sample.get("donations")
        for donation in donations:
            data_list.append({"person": person,
                "employed": employed,
                "location": location,
                "donation_type": donation.get("type"),
                "donation_date": donation.get("date"),
                "donation_value": donation.get("value")
                })

优化方案

方案1:统一循环逻辑,消除分支重复

核心思路是把空的donations列表替换成包含默认值的字典列表,无需分if/else分支,只用一套循环逻辑处理所有情况:

import pandas as pd

data_list = []
for sample in samples:
    # 提取通用基础字段,避免重复获取
    base_info = {
        "person": sample["person"],
        "employed": sample["employed"],
        "location": sample["location"]
    }
    # 处理捐赠数据:空列表则用默认值字典,否则保留原数据
    target_donations = sample["donations"] or [
        {"type": "NONE LISTED", "date": "NONE LISTED", "value": "NONE LISTED"}
    ]
    # 统一循环合并数据
    for donation in target_donations:
        data_list.append({
            **base_info,
            "donation_type": donation.get("type", "NONE LISTED"),
            "donation_date": donation.get("date", "NONE LISTED"),
            "donation_value": donation.get("value", "NONE LISTED")
        })

# 转换为DataFrame
df = pd.DataFrame(data_list)

方案2:利用Pandas内置API高效处理

如果熟悉Pandas,可以直接用explode和concat方法,完全避免手动循环,代码更简洁:

import pandas as pd

# 1. 先将原始数据转为DataFrame
df = pd.DataFrame(samples)
# 2. 展开donations列,空列表会转为NaN
df = df.explode("donations", ignore_index=True)
# 3. 替换NaN为默认字典
df["donations"] = df["donations"].apply(
    lambda x: x if x is not None else {"type": "NONE LISTED", "date": "NONE LISTED", "value": "NONE LISTED"}
)
# 4. 展开donations字典为单独列,并重命名匹配需求
donation_details = pd.DataFrame(df["donations"].tolist()).rename(
    columns={"type": "donation_type", "date": "donation_date", "value": "donation_value"}
)
# 5. 合并基础列和捐赠详情列
result_df = pd.concat([df.drop("donations", axis=1), donation_details], axis=1)

内容的提问来源于stack exchange,提问作者SMJune

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 22:45:37