如何将含层级字段的DataFrame转换为多级嵌套字典?
解决DataFrame转多级嵌套字典的问题
问题描述
现有包含country_code、customer_state、customer_city、returns_count列的DataFrame,示例数据如下:
[{'country_code': 'IN', 'customer_state': 'Uttar Pradesh', 'customer_city': 'Agra', 'returns_count': 100}, {'country_code': 'IN', 'customer_state': 'Uttar Pradesh', 'customer_city': 'Meerut', 'returns_count': 120}, {'country_code': 'IN', 'customer_state': 'Uttar Pradesh', 'customer_city': 'Lucknow', 'returns_count': 110}, {'country_code': 'IN', 'customer_state': 'Uttar Pradesh', 'customer_city': 'Noida', 'returns_count': 90}, {'country_code': 'IN', 'customer_state': 'Karnataka', 'customer_city': 'Bangalore', 'returns_count': 100}, {'country_code': 'IN', 'customer_state': 'Karnataka', 'customer_city': 'Mysore', 'returns_count': 200}, {'country_code': 'US', 'customer_state': 'California', 'customer_city': 'LA', 'returns_count': 180}, {'country_code': 'US', 'customer_state': 'California', 'customer_city': 'San Jose', 'returns_count': 150}, {'country_code': 'US', 'customer_state': 'California', 'customer_city': 'San Francisco', 'returns_count': 200}, {'country_code': 'US', 'customer_state': 'California', 'customer_city': 'San Diego', 'returns_count': 140}]
需要将其转换为三级嵌套字典:
- 一级键:
country_code - 二级键:
customer_state - 三级键:
customer_city - 对应值:包含
returns_count的字典
预期输出示例:
{'IN': {'Uttar Pradesh' : {'Agra' : {'returns_count':100}, 'Meerut' : {'returns_count':120}, 'Lucknow' : {'returns_count':110}, 'Noida' : {'returns_count' :90}}, 'Karnataka' : {'Bangalore' :{'returns_count':100}, 'Mysore' : {'returns_count' :200}}}, 'US': {'California' : {'LA' : {'returns_count':180}, 'San Jose' : {'returns_count':150}, 'San Francisco' : {'returns_count':200}, 'San Diego' : {'returns_count':140}}}}
用户尝试的代码报错:
df = df.groupby('country_code')[['customer_state', 'customer_city', 'returns_value', 'returns_count', 'orders_count', 'return_rate', 'latitude', 'longitude']].apply(lambda x:x.set_index('customer_state').to_dict(orient='index')).to_dict()
问题分析
原代码的问题在于:
- 仅按
country_code分组后直接转字典,无法实现三级嵌套结构,缺少对customer_state的二次分组处理 - 代码中引用了
returns_value等未在示例数据中出现的列,若DataFrame中不存在这些列会直接报错
正确实现方法
方法1:逐层分组构建嵌套字典
利用groupby逐层处理,先按国家分组,再对每组按州分组,最后将城市和对应值转为字典:
result = {} # 按国家分组 for country, country_group in df.groupby('country_code'): state_dict = {} # 按州分组 for state, state_group in country_group.groupby('customer_state'): # 将城市作为键,returns_count包装为字典作为值 city_dict = state_group.set_index('customer_city')['returns_count'].apply(lambda x: {'returns_count': x}).to_dict() state_dict[state] = city_dict result[country] = state_dict
方法2:使用层级索引+循环调整结构
先设置三级索引,再通过遍历调整为目标嵌套结构:
# 设置多级索引 df_indexed = df.set_index(['country_code', 'customer_state', 'customer_city']) # 初始化结果字典 result = {} # 遍历索引和值构建嵌套结构 for (country, state, city), count in df_indexed['returns_count'].items(): if country not in result: result[country] = {} if state not in result[country]: result[country][state] = {} result[country][state][city] = {'returns_count': count}
方法3:嵌套groupby+apply链式处理
通过嵌套的apply调用实现层级结构的自动构建:
result = df.groupby('country_code').apply( lambda country_group: country_group.groupby('customer_state').apply( lambda state_group: state_group.set_index('customer_city')['returns_count'].apply( lambda x: {'returns_count': x} ).to_dict() ).to_dict() ).to_dict()
以上三种方法均能生成符合预期的三级嵌套字典结构。
内容的提问来源于stack exchange,提问作者Pooja Bhateley
相关产品推荐
相关产品推荐

