如何基于另一个DataFrame的列对含重复行的DataFrame排序?
Got it, let's work through this problem. You're trying to sort df2 using the order defined in df1's Order column, but hitting a ValueError: cannot reindex from a duplicate axis when using .index() or .reindex()—that makes sense because df2 has duplicate entries for each name, and those methods rely on unique indexes to work properly.
Here's a straightforward solution that gets you exactly the sorted df you want:
Step-by-Step Solution
Create a mapping from Name to Order
First, we'll turn df1'sNameandOrdercolumns into a dictionary so we can easily assign the correct order value to each row in df2:order_map = df1.set_index('Name')['Order'].to_dict()This gives us a dict like
{'John':2, 'Alice':3, 'Alisha':1, ..., 'Steve':4}.Add an Order column to df2
Use the mapping to add anOrdercolumn to df2—this lets us tie each row in df2 to its corresponding priority from df1:df2['Order'] = df2['Name'].map(order_map)Now every row in df2 has the same
Ordervalue as its matching name in df1.Sort df2 by the Order column
Finally, sort df2 using the newOrdercolumn, then clean up by dropping the temporaryOrdercolumn and resetting the index (optional but nice for clean output):df_sort = df2.sort_values('Order').drop('Order', axis=1).reset_index(drop=True)
Why This Works
Unlike .reindex() which struggles with duplicate rows in df2, this method uses a numerical sort key (the Order column) that works seamlessly even with repeated names. All rows for the same name will stay grouped together, and the groups will be ordered exactly as defined in df1.
Final Result
Running this code will give you the exact df_sort you're expecting:
Name Condition Action 0 Alisha Stable Out 1 Alisha Unstable In 2 John Stable Out 3 John Unstable In 4 Alice Stable Out 5 Alice Unstable In 6 Steve Stable Out 7 Steve Unstable In 8 Mike Stable Out 9 Mike Unstable In 10 Katie Stable Out 11 Katie Unstable In
内容的提问来源于stack exchange,提问作者John Doe

