Pandas分组计算列差总和再求组内平均值的代码修正问题
Hey there! Let's fix this issue step by step. Your original code using apply isn't working because it's trying to process rows individually instead of grouping first, and the logic inside the lambda doesn't align with your requirement.
Let's Break Down Your Requirement First
You need to:
- Group the DataFrame by
yearandcode - For each group, calculate the average of (col2 - col1) (which is equivalent to summing all
col2-col1values in the group then dividing by the number of rows in the group) - Attach this average value to every row in its corresponding group
Original DataFrame (for reference)
First, let's confirm your input data:
import pandas as pd data = { 'year': [2019, 2019, 2019, 2018, 2018], 'code': [1, 1, 1, 2, 2], 'col1': [2, 3, 2, 1, 2], 'col2': [3, 5, 4, 4, 6] } df = pd.DataFrame(data)
Corrected Code
The simplest and most efficient way to achieve this is using groupby combined with transform — transform automatically broadcasts the group-level calculation result back to every row in the group:
# Calculate (col2 - col1) for all rows first, then group and get mean per group df['avg_num'] = (df['col2'] - df['col1']).groupby([df['year'], df['code']]).transform('mean')
Alternatively, if you prefer using groupby.apply (for more explicit group-level logic), this works too:
# Calculate mean of (col2-col1) per group, then map back to original rows grouped_mean = df.groupby(['year', 'code']).apply(lambda group: (group['col2'] - group['col1']).mean()) df['avg_num'] = df.set_index(['year', 'code']).index.map(grouped_mean)
Result You'll Get
After running either code, your DataFrame will look like this:
| year | code | col1 | col2 | avg_num |
|---|---|---|---|---|
| 2019 | 1 | 2 | 3 | 1.333... |
| 2019 | 1 | 3 | 5 | 1.333... |
| 2019 | 1 | 2 | 4 | 1.333... |
| 2018 | 2 | 1 | 4 | 3.5 |
| 2018 | 2 | 2 | 6 | 3.5 |
(Calculation check: For 2019/code1, (1+2+2)/3 = 5/3 ≈1.333; for 2018/code2, (3+4)/2=3.5)
Why Your Original Code Failed
Your original line df.apply(lambda row: (row['col_2'] - row['col_1']).mean(level=[0, 1]).reset_index(name='avg_num')) has two critical issues:
df.applyprocesses individual rows, sorow['col_2'] - row['col_1']is just a single number (scalar), not a Series. Scalars don't have a.mean()method, let alone alevelparameter.- The logic is backwards: you need to group first, then compute the mean per group — not process rows one by one.
内容的提问来源于stack exchange,提问作者daiyue

