Python3.7+Pandas1.1.3下如何将DataFrame转换为指定结构的嵌套字典?(去除冗余键并实现人口排序)
Convert Pandas DataFrame to Nested Dictionary (No Redundant Keys + Population Sort)
Let's fix your code to get the exact nested dictionary structure you want, while also handling the population-based sorting requirement for duplicate city names.
Corrected Code
import pandas as pd # Your original DataFrame setup location = {'city_id': [22000,25000,27000,35000], 'population': [3971883,2720546,667137,1323], 'region_name': ['California','Illinois','Massachusetts','Georgia'], 'city_name': ['Los Angeles','Chicago','Boston','Boston']} df = pd.DataFrame(location, columns = ['city_id', 'population','region_name', 'city_name']) # Generate the target nested dictionary result = df.sort_values(['city_name', 'population'], ascending=[True, False]) \ .groupby('city_name')[['region_name', 'city_id']] \ .apply(lambda x: x.set_index('region_name')['city_id'].to_dict()) \ .to_dict() print(result)
Output
{ 'Boston': {'Massachusetts': 27000, 'Georgia': 35000}, 'Chicago': {'Illinois': 25000}, 'Los Angeles': {'California': 22000} }
Key Fixes & Explanations
Remove Redundant "city_id" Key
Your original code usedx.set_index('region_name').to_dict()which converts the entire grouped sub-DataFrame (with thecity_idcolumn) into a dictionary. By adding['city_id']before.to_dict(), we directly convert only thecity_idvalues into a dictionary indexed byregion_name, eliminating the extra nestedcity_idkey.Sort by Population for Duplicate Cities
Thesort_values(['city_name', 'population'], ascending=[True, False])step ensures:- Groups are ordered by city name (optional, but keeps results clean)
- Within each city group, rows are sorted descending by population. Since Python 3.7+ preserves dictionary insertion order, this ensures the region with the larger population appears first in the nested dictionary for duplicate cities like Boston.
内容的提问来源于stack exchange,提问作者che kiria
相关产品推荐
相关产品推荐

