用Pythonic的DRY方式从嵌套列表与字典构建DataFrame
如何用更Pythonic且符合DRY原则的方式扁平化嵌套字典列表为DataFrame?
我有一个包含字典的列表,部分字典中还嵌套了字典列表,具体数据如下:
samples = [ {"person": "A", "employed": True, "location": "East", "donations": []}, {"person": "B", "employed": False, "location": "West", "donations": [ {"type": "cash", "date": "july 25, 2022", "value": 50}, {"type": "stocks", "date": "May 1, 2022", "value": 2500} ]}, {"person": "C", "employed": False, "location": "North", "donations": []} ]
我需要将这些数据处理为扁平化嵌套字典的DataFrame,使其具有指定的列头。我现有的代码存在大量重复,请问如何采用更Pythonic且符合DRY原则的实现方式?
现有代码如下:
data_list = [] ## initial empty list for sample in samples: ## get data that is not nested person = sample.get("person") employed = sample.get("employed") location = sample.get("location") ## if there is not nested data, give specific values if len(sample.get("donations")) < 1: donation_type = donation_date = donation_value = "NONE LISTED" ## build list data_list.append({"person": person, "employed": employed, "location": location, "donation_type": donation_type, "donation_date": donation_date, "donation_value": donation_value, }) else: ## build same list but with provided values donations = sample.get("donations") for donation in donations: data_list.append({"person": person, "employed": employed, "location": location, "donation_type": donation.get("type"), "donation_date": donation.get("date"), "donation_value": donation.get("value") })
优化方案
方案1:统一循环逻辑,消除分支重复
核心思路是把空的donations列表替换成包含默认值的字典列表,无需分if/else分支,只用一套循环逻辑处理所有情况:
import pandas as pd data_list = [] for sample in samples: # 提取通用基础字段,避免重复获取 base_info = { "person": sample["person"], "employed": sample["employed"], "location": sample["location"] } # 处理捐赠数据:空列表则用默认值字典,否则保留原数据 target_donations = sample["donations"] or [ {"type": "NONE LISTED", "date": "NONE LISTED", "value": "NONE LISTED"} ] # 统一循环合并数据 for donation in target_donations: data_list.append({ **base_info, "donation_type": donation.get("type", "NONE LISTED"), "donation_date": donation.get("date", "NONE LISTED"), "donation_value": donation.get("value", "NONE LISTED") }) # 转换为DataFrame df = pd.DataFrame(data_list)
方案2:利用Pandas内置API高效处理
如果熟悉Pandas,可以直接用explode和concat方法,完全避免手动循环,代码更简洁:
import pandas as pd # 1. 先将原始数据转为DataFrame df = pd.DataFrame(samples) # 2. 展开donations列,空列表会转为NaN df = df.explode("donations", ignore_index=True) # 3. 替换NaN为默认字典 df["donations"] = df["donations"].apply( lambda x: x if x is not None else {"type": "NONE LISTED", "date": "NONE LISTED", "value": "NONE LISTED"} ) # 4. 展开donations字典为单独列,并重命名匹配需求 donation_details = pd.DataFrame(df["donations"].tolist()).rename( columns={"type": "donation_type", "date": "donation_date", "value": "donation_value"} ) # 5. 合并基础列和捐赠详情列 result_df = pd.concat([df.drop("donations", axis=1), donation_details], axis=1)
内容的提问来源于stack exchange,提问作者SMJune
相关产品推荐
相关产品推荐

