You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas RollingGroupby多函数聚合时原索引丢失问题问询

Pandas RollingGroupby 聚合时索引不一致问题

当在Pandas中将groupby与rolling链式调用生成pandas.core.window.rolling.RollingGroupby对象后,使用agg函数会出现索引表现不一致的情况:

1. 单个聚合函数的正常表现

使用单个聚合函数(字符串形式)时,结果索引符合预期,同时包含分组标签和原DataFrame的整数索引:

import pandas as pd
data = pd.DataFrame({"data": [1,2,3,4,5]})
time = pd.date_range(start="1-1-2023", end="1-5-2023")
grp = pd.Series(["A", "A", "B", "B", "A"])
roll_gb = data.groupby(by=grp).rolling(window="2D", center=False, min_periods=1, on=time)

print(roll_gb.agg("mean"))

输出结果:

data
A 0   1.0
  1   1.5
  4   5.0
B 2   3.0
  3   3.5

2. 多聚合函数(含单函数列表)的异常表现

当使用多个聚合函数,哪怕是只包含一个函数的列表形式时,原整数索引会被替换为rolling调用中on参数指定的时间索引:

print(roll_gb.agg(["mean"]))

输出结果:

data
             mean
A 2023-01-01  1.0
  2023-01-02  1.5
  2023-01-05  5.0
B 2023-01-03  3.0
  2023-01-04  3.5

解决办法

如果需要在多聚合函数场景下保留原整数索引,可以在聚合后手动将时间索引映射回原索引:

# 先获取聚合结果
result = roll_gb.agg(["mean"])
# 创建原索引与时间的映射字典
idx_map = pd.Series(time.index, index=time)
# 替换内层索引
result.index = result.index.set_levels(idx_map[result.index.get_level_values(1)], level=1)
print(result)

输出结果:

data
         mean
A 0      1.0
  1      1.5
  4      5.0
B 2      3.0
  3      3.5

或者,也可以避免使用on参数,转而将时间列设置为DataFrame的索引后再进行groupby和rolling操作,这样聚合后的索引行为会更一致:

data = pd.DataFrame({"data": [1,2,3,4,5], "time": time, "grp": grp}).set_index("time")
roll_gb = data.groupby(by="grp")["data"].rolling(window="2D", center=False, min_periods=1)
# 单函数聚合
print(roll_gb.agg("mean"))
# 多函数聚合
print(roll_gb.agg(["mean"]))

两种方式的结果都会保留原时间索引,若需要原整数索引,可以额外添加到DataFrame中作为列参与后续映射。


内容的提问来源于stack exchange,提问作者LogZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 12:05:36