如何将Pandas DataFrame转指定结构字典?求Pythonic实现
更Pythonic的Pandas DataFrame转嵌套字典方法
需求说明
需要将给定的Pandas DataFrame转换为以Tract_number为顶级键,内部以统计指标为键、年份-数值对为值的嵌套字典结构。
原始DataFrame示例
import pandas as pd t = {'Tract_number': ['01001020100', '01001020100', '01001020100', '01001020100', '01001020100', '01001020100', '01001020100', '01001020100', '01001020100', '01001020100', '01001020100', '01001020100'], 'Year': [2019, 2014, 2015, 2016, 2017, 2018, 2011, 2020, 2010, 2009, 2012, 2013], 'Median_household_income': [70625.0, 65800.0, 67356.0, 68750.0, 70486.0, 70385.0, 66953.0, 70257.0, 71278.0, 'nan', 65179.0, 65114.0], 'Total_Asian_Population': [2.0, 12.0, 12.0, 9.0, 22.0, 17.0, 0.0, 41.0, 0.0, 'nan', 0.0, 0.0], 'Total_bachelors_degree': [205.0, 173.0, 166.0, 216.0, 261.0, 236.0, 139.0, 'nan', 170.0, 'nan', 156.0, 183.0], 'Total_graduate_or_professional_degree': [154.0, 149.0, 176.0, 191.0, 215.0, 174.0, 117.0, 'nan', 146.0, 'nan', 131.0, 127.0], 'Median_gross_rent': [749.0, 738.0, 719.0, 484.0, 780.0, 827.0, 398.0, 820.0, 680.0, 'nan', 502.0, 525.0]} df_sample = pd.DataFrame(data=t)
目标嵌套字典结构
A = { '01001020100': { 'Median_household_income': {'2010': 71278.0, '2011': 66953.0, ...}, 'Total_Asian_Population': {'2010': 0.0, '2011': 0.0, ...}, ... } }
用户现有实现
d = {'Tract_number': df_sample['Tract_number'].iloc[0]} e = { 'Median_household_income': pd.Series(df_sample.Median_household_income.values,index=df_sample.Year).to_dict(), 'Total_Asian_Population': pd.Series(df_sample.Total_Asian_Population.values,index=df_sample.Year).to_dict(), 'Total_bachelors_degree': pd.Series(df_sample.Total_bachelors_degree.values,index=df_sample.Year).to_dict(), 'Total_graduate_or_professional_degree': pd.Series(df_sample.Total_bachelors_degree.values,index=df_sample.Year).to_dict(), 'Median_gross_rent': pd.Series(df_sample.Total_bachelors_degree.values,index=df_sample.Year).to_dict() } f = {} f[d['Tract_number']] = e f
注意:现有代码存在笔误,
Total_graduate_or_professional_degree和Median_gross_rent错误引用了Total_bachelors_degree的值,需要修正。
更Pythonic的实现方法
方案1:针对单个Tract_number场景
利用字典推导式和Pandas内置方法,减少冗余代码,同时自动适配列名变化:
import pandas as pd # 先将字符串'nan'转换为Pandas可识别的缺失值(可选,按需保留或移除) df_sample = df_sample.replace('nan', pd.NA) # 生成目标嵌套字典 tract_id = df_sample['Tract_number'].iloc[0] output = { tract_id: { col: df_sample.set_index('Year')[col].to_dict() # 若要移除缺失值,可改为:df_sample.set_index('Year')[col].dropna().to_dict() for col in df_sample.columns.difference(['Tract_number', 'Year']) } }
方案2:支持多个Tract_number场景
如果DataFrame包含多个不同的Tract_number,可结合groupby实现批量转换:
import pandas as pd df_sample = df_sample.replace('nan', pd.NA) output = {} for tract, group in df_sample.groupby('Tract_number'): output[tract] = { col: group.set_index('Year')[col].to_dict() for col in group.columns.difference(['Tract_number', 'Year']) }
优势说明
- 简洁性:用字典推导式替代重复的
pd.Series创建代码,减少冗余 - 扩展性:通过
columns.difference自动排除不需要的列,新增统计指标时无需修改代码 - 健壮性:统一处理缺失值,避免字典中出现字符串'nan'的无效值
- 可读性:逻辑清晰,符合Python"简洁直观"的风格
内容的提问来源于stack exchange,提问作者Wolfy
相关产品推荐
相关产品推荐

