使用Pandas链式方法筛选组内行,排除两类创建量均为0的用户
Got it, let's solve this problem using pandas' chained methods as you requested. The goal is to exclude users where all their rows (both is_manually=True and is_manually=False) have created_per_week=0—like user_id 50 in your example.
Step 1: Recreate your sample DataFrame
First, let's set up the data to work with:
import pandas as pd df = pd.DataFrame({ 'user_id': [10, 10, 33, 33, 50, 50], 'is_manually': [True, False, True, False, True, False], 'created_per_week': [59, 90, 0, 64, 0, 0] })
Step 2: Chained Groupby Filter Method
The cleanest way to do this is using groupby().filter() in a chain. This method lets you apply a condition to each user group and keep only the groups that meet your criteria:
filtered_df = df.groupby('user_id').filter(lambda group: not (group['created_per_week'] == 0).all())
How this works:
groupby('user_id'): Groups the DataFrame by each unique user.lambda group: not (group['created_per_week'] == 0).all(): For each group, we check if all values increated_per_weekare 0. We negate this condition (not) so we keep groups where at least one row has a non-zero value.
Step 3: Alternative Chained Method with Transform
If you prefer using a mask to filter directly on the original DataFrame, you can use transform() in a chain:
filtered_df = df[df.groupby('user_id')['created_per_week'].transform(lambda x: not (x == 0).all())]
How this works:
transform()returns a boolean array with the same length as the original DataFrame, marking which rows belong to groups that aren't all zeros.- We use this array to index the original DataFrame and keep only the relevant rows.
Result
Both methods will give you this filtered DataFrame:
user_id is_manually created_per_week 0 10 True 59 1 10 False 90 2 33 True 0 3 33 False 64
Notice user_id 50 is excluded, as both their rows had created_per_week=0.
内容的提问来源于stack exchange,提问作者ytu

