如何基于指定列表对Pandas DataFrame的行进行自定义排序?
Got it, let's solve this problem where you need to sort your DataFrame rows to match the exact order of a given list. Using your example data:
import pandas as pd df = pd.DataFrame({'user': ['Bob', 'Jane', 'Alice'], 'income': [40000, 50000, 42000]}) z = ["Jane", "Alice", "Bob"]
We have two straightforward methods to get the desired sorted DataFrame:
Method 1: Using Categorical Data Type
This method is great because it explicitly defines the order you want, and works even if there are duplicates in your DataFrame (though your problem states each value appears once). Here's how:
- Convert the
usercolumn to a Categorical type, specifying your custom list as the ordered categories:
df['user'] = pd.Categorical(df['user'], categories=z, ordered=True)
- Now sort the DataFrame by the
usercolumn—pandas will respect the custom order you set:
df_sorted = df.sort_values('user')
- (Optional) If you don't want to keep the
usercolumn as a Categorical type, convert it back to a regular string:
df_sorted['user'] = df_sorted['user'].astype(str)
The resulting df_sorted will look like this:
| user | income |
|---|---|
| Jane | 50000 |
| Alice | 42000 |
| Bob | 40000 |
Method 2: Using set_index + reindex
Since your list z contains exactly the values from the user column (each once), you can reindex the DataFrame directly:
- Set the
usercolumn as the index of the DataFrame:
df_indexed = df.set_index('user')
- Reindex the DataFrame using your custom list
z—this will reorder the rows to match the list:
df_sorted = df_indexed.reindex(z).reset_index()
This achieves the same result as Method 1, and is a bit more concise if you know all values in z exist in the DataFrame.
Both methods work perfectly for your use case—pick whichever feels more intuitive to you!
内容的提问来源于stack exchange,提问作者user3680510

