如何在Pandas多层索引DataFrame中保留Column2为xxx的行
Solution to Filter Your Entry-Based DataFrame
First, let's address the structure of your DataFrame—those blank values in Column1 are just display shorthand for rows belonging to the same Entry. We'll start by fixing that to ensure every row is properly linked to its parent Entry, then apply the filter you need.
Step 1: Clean Up the Entry Column
First, let's recreate your data and fill in the missing Entry values:
import pandas as pd # Your original data structure data = { 'Column1': ['Entry1', '', '', 'Entry2', '', '', 'Entry3', '', ''], 'Column2': ['xxx', 'ggg', 'hhh', 'xxx', 'ggg', 'hhh', 'xxx', 'ggg', 'xxx'], 'Column3': ['yyyy', 'ffff', 'llll', 'yyyy', 'ffff', 'llll', 'yyyy', 'ffff', 'llll'], 'Column4': [3151, 4647, 46847, 215, 12321, 1321, 4564, 6654, 46542], 'Column5': [6542, 587, 321321, 1231, 5487, 5854, 123, 135, 1348] } df = pd.DataFrame(data) # Fill blank Column1 values with the last valid Entry (forward fill) df['Column1'] = df['Column1'].ffill()
Step 2: Filter for Rows Where Column2 is 'xxx'
Now we'll narrow down to only the rows you want, keeping just the first occurrence of 'xxx' per Entry:
# Keep only rows where Column2 equals 'xxx' filtered_rows = df[df['Column2'] == 'xxx'] # Retain just the first 'xxx' row for each Entry final_result = filtered_rows.drop_duplicates(subset='Column1', keep='first') # Alternative method using groupby (same end result): # final_result = filtered_rows.groupby('Column1').head(1)
Step 3: Check the Output
Printing final_result will give you exactly the output you expected:
Column1 Column2 Column3 Column4 Column5 0 Entry1 xxx yyyyy 3151 6542 3 Entry2 xxx yyyyy 215 1231 6 Entry3 xxx yyyyy 4564 123
Quick Notes:
- If your original DataFrame already has Column1 fully populated (no blanks), skip the forward fill step.
- If you ever want to keep all 'xxx' rows for an Entry (not just the first), just use
filtered_rowsdirectly without the duplicate drop/groupby step.
内容的提问来源于stack exchange,提问作者Crovish
相关产品推荐
相关产品推荐

