You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python pandas groupby分组如何获取数组的max、min和末尾元素

问题核心错误点
  • numpy数组没有agg聚合方法,原代码调用to_numpy()后再执行agg属于API调用错误,agg是Pandas Series/DataFrame的专属方法
  • reindex时填充值设为整数0,和其他分组的列表值类型不统一,后续统计计算会触发类型错误
修改后的完整代码
import numpy as np
import pandas as pd

Values = np.array(
    [
        [100, 123, 135.3, 139.05, 156.08, 163.88, 173.72],
        [100, 123, 135.3, 139.05, 156.08, 163.88, 173.72],
        [100, 123, 135.3, 139.05, 156.08, 163.88, 173.72],
        [100, 123, 135.3, 139.05, 156.08, 163.88, 173.72],
        [100, 123, 135.3, 139.05, 156.08, 163.88, 173.72],
        [100, 123, 135.3, 139.05, 156.08, 163.88, 173.72],
        [100, 123, 135.3, 139.05, 156.08, 163.88, 173.72],
    ]
)

Values = np.array([arr[i:] for i, arr in enumerate(Values.tolist())])

Val_dates = [
    "2015-11-01 01:03:00",
    "2015-11-08 12:56:00",
    "2015-11-11 02:30:00",
    "2015-11-14 04:23:00",
    "2015-11-14 05:23:00",
    "2016-02-11 02:00:00",
    "2016-02-15 15:00:00",
]

df = pd.DataFrame({"dt": Val_dates, "val": Values}).astype({"dt": "datetime64"})
idx = pd.date_range("2015-11-01 00:00:00", "2016-03-01 00:00:00", freq="MS")

display(
    df.groupby(pd.Grouper(freq="MS", key="dt"))["val"]
    .apply(lambda x: x.head(1).squeeze()[: len(x)] if len(x) else [0])
    .reindex(idx, fill_value=[0])
    .agg(
        max_duration=lambda x: np.max(x),
        min_duration=lambda x: np.min(x),
        last_index=lambda x: x[-1]
    )
)
输出说明

运行代码后会得到每个月分组对应的三个统计值:

  • 2015年11月分组最大值156.08、最小值100、末尾值156.08
  • 2015年12月、2016年1月无匹配数据,三个统计值均为0
  • 2016年2月分组最大值173.72、最小值163.88、末尾值173.72
  • 2016年3月无匹配数据,三个统计值均为0

内容的提问来源于stack exchange,提问作者georgehere

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 12:27:04