Pandas筛选DataFrame中仅包含3和4的指定分组
Got it, let's work through this problem together. You need to keep only the groups that have both 3 and 4 in col1, with no other values at all—so groups B and D are the ones we want to retain from your sample data.
Step 1: Set Up the Sample Data
First, let's create the DataFrame with your example values so we can test our solution directly:
import pandas as pd data = { 'groups': ['A', 'A', 'A', 'A', 'B', 'B', 'B', 'C', 'D', 'D'], 'col1': [3, 4, 2, 1, 3, 3, 4, 2, 4, 3] } df = pd.DataFrame(data)
Step 2: Filter Valid Groups
There are a couple of clean ways to do this. Let's start with the most efficient one:
Method 1: Use Groupby + Set Aggregation
This approach leverages set operations to check if a group's values exactly match {3,4}:
# Get the set of unique values for each group group_values = df.groupby('groups')['col1'].agg(set) # Filter groups where the set is exactly {3,4} valid_groups = group_values[group_values == {3, 4}].index # Keep only rows from valid groups in the original DataFrame result = df[df['groups'].isin(valid_groups)]
Method 2: Use Groupby Filter with Lambda (More Explicit)
If you prefer a more verbose, readable approach, you can use filter() with a lambda that checks three clear conditions:
- The group contains 3
- The group contains 4
- All values in the group are either 3 or 4
valid_groups = df.groupby('groups').filter( lambda x: (3 in x['col1'].values) and (4 in x['col1'].values) and (x['col1'].isin([3, 4]).all()) ).groups.keys() result = df[df['groups'].isin(valid_groups)]
Step 3: Check the Result
Either method will give you the expected output:
groups col1 4 B 3 5 B 3 6 B 4 8 D 4 9 D 3
Both methods work, but Method 1 is faster for large datasets since set aggregation is more efficient than row-wise checks in the lambda.
内容的提问来源于stack exchange,提问作者chippycentra

