You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何对带多重索引的Pandas DataFrame按多列分组求FFDI_SFC最大值

实现方案

基础实现(兼容绝大多数Pandas版本)

直接通过groupby混合指定索引级别和列名进行分组计算,代码如下:

# 分组计算日度最大FFDI_SFC,结果默认保留分组键为多重索引
result = df.groupby(
    ["latitude", "longitude", "AET_date"],
    as_index=True
)["FFDI_SFC"].max().to_frame(name="daily_max_FFDI_SFC")

注:计算过程中Pandas会自动忽略FFDI_SFC的空值,若某分组内所有FFDI_SFC均为空,对应结果返回NaN,符合可空字段的要求

大数量级性能优化

你的数据集规模超过1亿行,可根据实际情况选择以下优化手段:

  • 内存优化:提前仅保留计算所需字段,减少中间数据内存占用
    temp_df = df[["FFDI_SFC", "AET_date"]].copy()
    result = temp_df.groupby(["latitude", "longitude", "AET_date"])["FFDI_SFC"].max().to_frame("daily_max_FFDI_SFC")
    
  • 空值优化:若FFDI_SFC空值占比高,先剔除空行再计算
    df = df.dropna(subset=["FFDI_SFC"])
    result = df.groupby(["latitude", "longitude", "AET_date"])["FFDI_SFC"].max().to_frame("daily_max_FFDI_SFC")
    
  • 分类键优化:若分组键为Categorical分类类型,添加observed=True避免生成不存在的分组组合
    result = df.groupby(
        ["latitude", "longitude", "AET_date"],
        as_index=True,
        observed=True
    )["FFDI_SFC"].max().to_frame(name="daily_max_FFDI_SFC")
    

正确性验证

可以先截取小批量数据验证逻辑是否符合预期:

# 取前1万行测试
test_result = df.head(10000).groupby(["latitude", "longitude", "AET_date"])["FFDI_SFC"].max().to_frame("daily_max_FFDI_SFC")
print(test_result.index.names) # 预期输出:['latitude', 'longitude', 'AET_date']
print(test_result.columns) # 预期输出:['daily_max_FFDI_SFC']

内容的提问来源于stack exchange,提问作者alextc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.26 18:24:05