如何筛选Pandas多层索引中外层组内Count均大于10的行?
Great question! Looping through groups can get slow and messy with large datasets, so leaning into Pandas' built-in vectorized operations is way cleaner and faster. Here are two optimal solutions to solve your problem:
First, let's recreate your sample data (I added a seed for reproducibility):
import numpy as np import pandas as pd np.random.seed(42) # Ensures consistent random values for testing data = pd.DataFrame(np.random.rand(8), index=[list('AABBCCDD'), ['M', 'F']*4]) data['Count'] = [1,2,15,17,8,12,11,20]
Approach 1: Using groupby().filter()
This method directly groups by the outer index and keeps only groups where all 'Count' values meet your threshold. Since each outer group has exactly two entries ('M' and 'F'), checking if all are >10 guarantees both conditions are satisfied:
# Filter groups where every 'Count' value is greater than 10 filtered_data = data.groupby(level=0).filter(lambda group: (group['Count'] > 10).all())
How it works:
groupby(level=0)groups the DataFrame by the outer index (A, B, C, D).- The
filter()method runs the lambda function on each group:(group['Count'] >10).all()returnsTrueonly if both 'M' and 'F' in the group have a Count >10. - Only groups that pass this check stay in the final result.
Approach 2: Unstack and Explicitly Check Conditions
If you want to be more explicit about validating 'M' and 'F' individually (handy if your groups might grow beyond two entries later), unstack the inner index, apply your conditions, then filter the original DataFrame:
# Unstack the inner index to turn 'M' and 'F' into columns unstacked = data.unstack(level=1) # Create a mask where both 'M' and 'F' have Count >10 valid_groups = (unstacked['Count']['M'] > 10) & (unstacked['Count']['F'] > 10) # Filter the original DataFrame using the valid group indices filtered_data = data.loc[valid_groups.index]
Output for both approaches:
Running either method will give you the desired result:
0 Count B M 0.020584 15 F 0.969910 17 D M 0.181825 11 F 0.183405 20
Both methods skip manual looping and use Pandas' optimized operations, making them far more efficient—especially with large datasets.
内容的提问来源于stack exchange,提问作者elz

