You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中按分组首条最大值过滤DataFrame?

简化Polars分组取最大值首次出现的方法

你可以通过排序后去重的方式简化实现,仅需指定分组列、最大值列,再配合用于确定"首次出现"的顺序列(如时间列或行号)即可,代码更简洁高效:

核心实现

import polars as pl

df = pl.DataFrame(
    {
        "cat": [1, 1, 1, 2, 2, 2, 2, 3, 3, 3],
        "max_col": [12, 24, 36, 15, 50, 50, 45, 20, 40, 60],
        "other_col": [25, 50, 75, 125, 150, 175, 200, 225, 250, 275],
    }
)

# 按分组列升序、最大值列降序、顺序列升序排序,再取每个分组的第一条记录
result = df.sort(
    ["cat", "max_col", "other_col"],
    descending=[False, True, False]
).unique("cat", keep="first")

print(result)

输出结果和你原代码一致:

shape: (3, 3)
┌─────┬─────────┬───────────┐
│ cat ┆ max_col ┆ other_col │
│ --- ┆ ---     ┆ ---       │
│ i64 ┆ i64     ┆ i64       │
╞═════╪═════════╪═══════════╡
│ 1   ┆ 36      ┆ 75        │
│ 2   ┆ 50      ┆ 150       │
│ 3   ┆ 60      ┆ 275       │
└─────┴─────────┴───────────┘

封装成通用函数

如果需要更灵活的复用(仅传入分组列和最大值列,可选指定顺序列),可以封装成函数:

def get_first_max_row(df, group_col, max_col, order_col=None):
    # 若未指定顺序列,用原数据的行号作为顺序依据(保证取原数据中首次出现的最大值)
    if order_col is None:
        df = df.with_row_index("row_order")
        order_col = "row_order"
    
    sorted_df = df.sort(
        [group_col, max_col, order_col],
        descending=[False, True, False]
    ).unique(group_col, keep="first")
    
    # 自动移除临时行号列
    if order_col == "row_order":
        sorted_df = sorted_df.drop("row_order")
    
    return sorted_df

# 使用示例:仅传入分组列和最大值列,自动按原数据顺序取首次出现的最大值
result = get_first_max_row(df, "cat", "max_col")

# 也可以指定顺序列(如时间列)
result = get_first_max_row(df, "cat", "max_col", "other_col")

方法优势

  1. 代码更简洁:无需嵌套两次filter,逻辑直观
  2. 性能更优:排序+去重的操作在Polars中经过高度优化,比多次窗口过滤效率更高
  3. 灵活性强:可直接指定顺序列,或默认按原数据行号取首次出现的记录

内容的提问来源于stack exchange,提问作者niko86

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 04:17:16