You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

pandas如何计算分组内前两个年度数据的平均值并存为新列

实现代码

提供两种实现方案,可按需选择:

方案1:滚动窗口实现(通用可扩展)

适合数据存在更多年份的场景,会自动对应每行的前两年计算均值,无需硬编码年份:

import pandas as pd

# 构造原始数据
d2 = {'year': ['1996', '1997', '1998', '1996',  '1997', '1998'], 
      'country': ['USA', 'USA', 'USA', 'ITA', 'ITA', 'ITA'],
      'value': [0, 1, 3, 2.75, 1, 8]}
df2 = pd.DataFrame(data=d2)

# (可选)如果数据年份无序,先按国家+年份升序排序保证计算准确
df2 = df2.sort_values(['country', 'year'], ignore_index=True)

# 分组计算滚动2年的均值,偏移1行赋值给当前行
df2['2_prior_year_mean'] = df2.groupby('country')['value']\
                              .rolling(window=2).mean()\
                              .shift(1)\
                              .reset_index(level=0, drop=True)

方案2:固定匹配指定年份(仅针对1996、1997年求均值赋值给1998年)

如果需求是固定只计算1996和1997年的均值,仅给1998年的行赋值,可以用这个更直接的写法:

import pandas as pd

# 构造原始数据
d2 = {'year': ['1996', '1997', '1998', '1996',  '1997', '1998'], 
      'country': ['USA', 'USA', 'USA', 'ITA', 'ITA', 'ITA'],
      'value': [0, 1, 3, 2.75, 1, 8]}
df2 = pd.DataFrame(data=d2)

# 计算每个国家1996、1997年的value均值
prior_mean = df2[df2['year'].isin(['1996','1997'])].groupby('country')['value'].mean()

# 仅给1998年的行赋值均值,其余行设为NaN
df2['2_prior_year_mean'] = df2.apply(
    lambda x: prior_mean[x['country']] if x['year'] == '1998' else pd.NA, 
    axis=1
)

运行结果

两种方案执行后均可得到预期输出:

year country  value  2_prior_year_mean
0  1996     USA   0.00                NaN
1  1997     USA   1.00                NaN
2  1998     USA   3.00              0.500
3  1996     ITA   2.75                NaN
4  1997     ITA   1.00                NaN
5  1998     ITA   8.00              1.875

内容的提问来源于stack exchange,提问作者user14237226

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.02 09:09:03