Pandas数据处理:当分组结果元素数大于1时删除对应零值行
Got it, let's tackle this problem step by step. You want to filter your original DataFrame based on grouped results: when a group (grouped by id_client and date) has multiple count values (forming a list after your apply(list) command), you need to remove rows where count is 0; if the group only has a single count value, keep that row no matter if it's 0 or not.
Method 1: Direct Group Filtering (Most Efficient)
You don't even need to generate the list_count intermediate DataFrame to achieve this. We can use groupby().filter() to handle the logic directly on the original data:
First, let's recreate your sample DataFrame for testing:
import pandas as pd # Original sample data data = { 'id_client': [908, 908, 907, 907, 909, 910], 'date': ['01/2020'] * 6, 'count': [0, 35, 0, 37, 50, 0] } df = pd.DataFrame(data)
Now apply the filtering logic:
# Filter groups based on their size and count values filtered_df = df.groupby(['id_client', 'date']).filter( lambda group: True if len(group) == 1 else group['count'] != 0 ) # Reset index (optional but cleaner) filtered_df = filtered_df.reset_index(drop=True) print(filtered_df)
Explanation:
groupby(['id_client', 'date']).filter()applies the lambda function to each group. Rows where the function returnsTrueare kept.- For groups with only 1 row (equivalent to your single-element grouping result), we keep all rows (
True). - For groups with multiple rows (equivalent to your list-based grouping result), we only keep rows where
count != 0.
Expected Output:
id_client date count 0 908 01/2020 35 1 907 01/2020 37 2 909 01/2020 50 3 910 01/2020 0
Method 2: Using Your Intermediate list_count DataFrame
If you want to work with the list_count DataFrame you generated, you can merge it back to the original data and apply filtering:
First, generate your list_count as you did:
list_count = df.groupby(['id_client', 'date'])['count'].apply(list).reset_index()
Then merge and filter:
# Merge group info into original DataFrame df_with_groups = df.merge( list_count.rename(columns={'count': 'group_values'}), on=['id_client', 'date'] ) # Apply filtering logic filtered_df = df_with_groups[ # Keep rows where group is a single value OR count is not 0 (df_with_groups['group_values'].apply(lambda x: not isinstance(x, list) or len(x) == 1)) | (df_with_groups['count'] != 0) ].drop(columns=['group_values']).reset_index(drop=True)
This will give you the same expected output as Method 1.
内容的提问来源于stack exchange,提问作者tamo007

