Python中使用预设排名对字符串变量频率分布排序失效问题排查
Hey there! I see the issue with your code—let's break it down and fix that custom sorting problem.
What's Going Wrong
Your line fd_education = fd_education.sort_index(level='education_rank') isn't doing what you expect. The level parameter in sort_index() is designed for multi-level indexes (when your index has multiple hierarchical levels), but your fd_education has a plain string index of education categories. So this line essentially does nothing, and sort_index() falls back to its default behavior: alphabetical sorting.
Two Working Solutions
Solution 1: Sort Using Your Custom Rank Map Directly
We can sort the index of your frequency series using the rank values from your education_rank dictionary, then reorder the series with that sorted index:
import pandas as pd # Your custom rank dictionary education_rank = {' Bachelors':12, ' HS-grad':8, ' 11th':6, ' Masters':14, ' 9th':5, ' Some-college':11, ' Assoc-acdm':10, ' Assoc-voc':9, ' 7th-8th':4, ' Doctorate':15, ' Prof-school':13, ' 5th-6th':3, ' 10th':16, ' 1st-4th':2, ' Preschool':1, ' 12th':7} # Calculate frequency counts fd_education = pd.value_counts(adult_data.education) # Sort the index by your custom rank, then reorder the series fd_education_sorted = fd_education.loc[sorted(fd_education.index, key=lambda x: education_rank[x])] print(fd_education_sorted)
Here, sorted(fd_education.index, key=lambda x: education_rank[x]) sorts the education categories based on their corresponding rank values, and loc[] reorders the frequency series to match this sorted list.
Solution 2: Convert to Categorical Data (More Robust)
If you plan to work with this sorted order repeatedly, converting the education column to an ordered categorical type is a better approach. This way, any subsequent operations (like value_counts) will respect your custom order automatically:
import pandas as pd # Your custom rank dictionary education_rank = {' Bachelors':12, ' HS-grad':8, ' 11th':6, ' Masters':14, ' 9th':5, ' Some-college':11, ' Assoc-acdm':10, ' Assoc-voc':9, ' 7th-8th':4, ' Doctorate':15, ' Prof-school':13, ' 5th-6th':3, ' 10th':16, ' 1st-4th':2, ' Preschool':1, ' 12th':7} # Get the education categories sorted by your rank (ascending order) sorted_education_order = sorted(education_rank.keys(), key=lambda x: education_rank[x]) # Convert the education column to an ordered categorical adult_data['education'] = pd.Categorical( adult_data['education'], categories=sorted_education_order, ordered=True ) # Now value_counts will output in your custom order fd_education = pd.value_counts(adult_data.education) print(fd_education)
This method makes the custom order part of the column's metadata, so you don't have to re-sort every time you calculate frequencies.
内容的提问来源于stack exchange,提问作者Dat

