如何在Pandas的apply(lambda)中集成count聚合函数并简化分组逻辑
Integrate Count Aggregation into Pandas
.apply(lambda x: ...) Got it, you’re looking to streamline your workflow by ditching the separate groupby count and merge steps—let’s get that count logic directly into your apply lambda!
The key insight here is that when you use groupby().apply(), each x in the lambda is the subset DataFrame for that group. The number of rows in this subset (which is exactly what groupby.size() gives you) is just len(x). We can add this as a new field right in the Series we’re building inside the lambda.
Here’s the simplified, one-step code that replaces your original three-part process:
import pandas as pd # Example DataFrame (replace with your actual data) df = pd.DataFrame({ 'id': [1, 1, 1, 2, 2], 'target': ['A', 'A', 'B', 'A', 'A'], 'duration': [10, 20, 15, 5, 25], 'status': ['active', 'inactive', 'active', 'active', 'inactive'], 'src': ['web', 'app', 'web', 'app', 'web'] }) # Combined aggregation with count included directly df_gp = df.groupby(['id', 'target']).apply(lambda x: pd.Series({ 'counts': len(x), # This replaces the separate groupby.size() step 'min_duration': x['duration'].min(), 'max_duration': x['duration'].max(), 'total_duration': x['duration'].sum(), 'all_status': list(x['status']), 'last_status': x['status'].iloc[-1], # More efficient than list indexing 'all_src': list(x['src']) })).reset_index() # Output the result print(df_gp)
Quick notes on improvements:
- Instead of
min(x['duration']), usingx['duration'].min()is more idiomatic Pandas and slightly more efficient. - For
last_status,x['status'].iloc[-1]avoids converting the entire column to a list just to grab the last element—better for performance with large datasets. - No more need for
df_countor thepd.merge()step; all aggregations happen in one pass over the grouped data.
内容的提问来源于stack exchange,提问作者Edamame
相关产品推荐
相关产品推荐

