按指定含重复值的列表顺序选取Pandas DataFrame行并保留原索引
Solution for Pandas DataFrame Row Selection with All Required Conditions
First, let's make sure we nail every requirement you listed:
- Return rows in the exact order of your specified value list
- Include duplicate rows for every duplicate entry in the list
- Keep the original DataFrame's index intact
- Ignore any values in the list that don't exist in your target column
Here's a straightforward, robust method that checks all these boxes, using your sample data to demonstrate:
Step-by-Step Code Implementation
import pandas as pd # Your sample DataFrame df = pd.DataFrame({'A': [5, 6, 3, 4], 'B': [1, 2, 3, 5]}) # Your target value list list_of_values = [3, 4, 6, 4, 3, 8] # 1. Map values in column 'A' to their original indices # This works even if a value appears multiple times in the column value_to_indices = df.reset_index().set_index('A')['index'] # 2. Filter out values from the list that aren't present in column 'A' filtered_values = [val for val in list_of_values if val in value_to_indices.index.unique()] # 3. Fetch rows in the filtered list order, including duplicates, with original indices result = df.loc[value_to_indices.reindex(filtered_values).values]
Result Verification
Running this code gives you exactly the output you're looking for:
A B 2 3 3 3 4 5 1 6 2 3 4 5 2 3 3
How This Meets Every Requirement
- Matches list order: We use
reindex(filtered_values)to strictly follow the order of your input list (after removing missing values) - Duplicate rows for duplicate values: The
reindexmethod preserves duplicate entries in the filtered list, sodf.loc[]returns the corresponding rows multiple times - Original index preserved: We're directly indexing the original DataFrame using its original index values, so no index is altered or replaced
- Ignores non-existent values: The list comprehension in step 2 weeds out any values that don't appear in column 'A' of your original DataFrame
Handling Duplicate Values in the Target Column
If your target column has duplicate values (e.g., two rows with A=3), the above method returns the first occurrence of each value by default. To return all matching rows for each value in the list (in order), adjust the mapping step like this:
# Map values to a list of all their corresponding indices value_to_indices_list = df.reset_index().groupby('A')['index'].apply(list).to_dict() # Expand the filtered list to include every matching index expanded_indices = [] for val in filtered_values: expanded_indices.extend(value_to_indices_list[val]) # Retrieve the full set of matching rows result = df.loc[expanded_indices]
For example, if your DataFrame had two rows with A=3:
df = pd.DataFrame({'A': [3, 6, 3, 4], 'B': [1, 2, 3, 5]}, index=[0, 1, 2, 3]) list_of_values = [3, 4]
The adjusted code would return:
A B 0 3 1 2 3 3 3 4 5
内容的提问来源于stack exchange,提问作者Lightspark
相关产品推荐
相关产品推荐

