You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas分组计算列差总和再求组内平均值的代码修正问题

Hey there! Let's fix this issue step by step. Your original code using apply isn't working because it's trying to process rows individually instead of grouping first, and the logic inside the lambda doesn't align with your requirement.

Let's Break Down Your Requirement First

You need to:

  1. Group the DataFrame by year and code
  2. For each group, calculate the average of (col2 - col1) (which is equivalent to summing all col2-col1 values in the group then dividing by the number of rows in the group)
  3. Attach this average value to every row in its corresponding group

Original DataFrame (for reference)

First, let's confirm your input data:

import pandas as pd

data = {
    'year': [2019, 2019, 2019, 2018, 2018],
    'code': [1, 1, 1, 2, 2],
    'col1': [2, 3, 2, 1, 2],
    'col2': [3, 5, 4, 4, 6]
}
df = pd.DataFrame(data)

Corrected Code

The simplest and most efficient way to achieve this is using groupby combined with transform — transform automatically broadcasts the group-level calculation result back to every row in the group:

# Calculate (col2 - col1) for all rows first, then group and get mean per group
df['avg_num'] = (df['col2'] - df['col1']).groupby([df['year'], df['code']]).transform('mean')

Alternatively, if you prefer using groupby.apply (for more explicit group-level logic), this works too:

# Calculate mean of (col2-col1) per group, then map back to original rows
grouped_mean = df.groupby(['year', 'code']).apply(lambda group: (group['col2'] - group['col1']).mean())
df['avg_num'] = df.set_index(['year', 'code']).index.map(grouped_mean)

Result You'll Get

After running either code, your DataFrame will look like this:

yearcodecol1col2avg_num
20191231.333...
20191351.333...
20191241.333...
20182143.5
20182263.5

(Calculation check: For 2019/code1, (1+2+2)/3 = 5/3 ≈1.333; for 2018/code2, (3+4)/2=3.5)

Why Your Original Code Failed

Your original line df.apply(lambda row: (row['col_2'] - row['col_1']).mean(level=[0, 1]).reset_index(name='avg_num')) has two critical issues:

  1. df.apply processes individual rows, so row['col_2'] - row['col_1'] is just a single number (scalar), not a Series. Scalars don't have a .mean() method, let alone a level parameter.
  2. The logic is backwards: you need to group first, then compute the mean per group — not process rows one by one.

内容的提问来源于stack exchange,提问作者daiyue

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 08:27:00