You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多索引时间序列DataFrame重采样报错ValueError的解决方法问询

解决方案

你的错误源于仅指定时间维度的Grouper,但原DataFrame使用的是(index, uid)的多索引结构,pd.Grouper无法直接处理这种多索引的分组逻辑。同时你的需求是按每个uid分别对时间序列做30分钟窗口聚合,需要同时将uid和时间窗口作为分组维度,以下是两种可行的实现方式:

方式一:同时按uid和时间窗口分组

直接通过groupby同时指定uid和时间Grouper(需明确时间对应的索引level):

result = df3.groupby([
    pd.Grouper(level='index', freq='30Min', closed='right', label='right'),
    pd.Grouper(level='uid')
]).agg({
    "col1": "max", 
    "col2": "min"
})

注:无需在agg中处理uid,因为已经按uid分组,结果会自动保留uid作为索引的一部分。

方式二:先按uid拆分,再重采样

先按uid分组拆分数据,再对每个用户的时间序列单独做重采样:

# 按uid分组后重采样
result = df3.groupby('uid').resample(
    '30Min', closed='right', label='right', level='index'
).agg({
    "col1": "max", 
    "col2": "min"
})
# 可选:调整索引顺序,让时间在前、uid在后
result = result.swaplevel('uid', 'index').sort_index()

额外提示

原生成数据代码中的np.random.random_integers已被numpy弃用,建议替换为np.random.randint以兼容新版本:

# 替换原代码中的两行
df['col2'] = np.random.randint(1, 101, size=df.shape[0])
df2['col2'] = np.random.randint(1, 51, size=df2.shape[0])

内容的提问来源于stack exchange,提问作者prof32

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 20:11:46