优化Pandas GroupBy+Apply分组均值计算的高效实现方案
When dealing with 73,800 groups, using apply() is going to be slow—since it runs custom Python logic for every single group, which adds a ton of overhead. Let's switch to vectorized pandas operations that leverage optimized C-backed code to cut down your runtime drastically.
Option 1: Use transform() + duplicated()
This approach keeps your original DataFrame structure while efficiently computing group means and filtering to the first row of each group:
# Step 1: Compute group-wise mean of 'c' and attach it to every row in the group df['c_mean'] = df.groupby(['a', 'b'])['c'].transform('mean') # Step 2: Keep only the first row of each group filtered = df[~df.duplicated(subset=['a', 'b'])].copy() # Step 3: Replace 'c' with the group mean, then clean up the temporary column filtered['c'] = filtered['c_mean'] final = filtered.drop('c_mean', axis=1).reset_index(drop=True)
Option 2: Use groupby.agg() (Most Efficient for This Use Case)
This method directly aggregates the data in one pass, avoiding any temporary columns. You just need to specify how to handle each column:
- For 'c': take the group mean
- For all other columns (like 'd'): take the first value from the group
final = df.groupby(['a', 'b']).agg( c=('c', 'mean'), d=('d', 'first') # Add other columns here with 'first' as needed ).reset_index()
Why These Methods Are Faster
Unlike apply(), which loops through each group and runs your custom Python function:
transform()andagg()are vectorized operations executed in pandas' optimized C backend.- They avoid the overhead of launching a Python function for every single group (73k times in your case!).
Example Output
For your sample DataFrame, you'll get exactly the result you need:
| a | b | c | d |
|---|---|---|---|
| "f" | "e" | 2.5 | True |
| "c" | "a" | 1.0 | True |
Expect this to run in seconds instead of minutes for your dataset.
内容的提问来源于stack exchange,提问作者LizzAlice

