如何在Polars中按分组首条最大值过滤DataFrame?
简化Polars分组取最大值首次出现的方法
你可以通过排序后去重的方式简化实现,仅需指定分组列、最大值列,再配合用于确定"首次出现"的顺序列(如时间列或行号)即可,代码更简洁高效:
核心实现
import polars as pl df = pl.DataFrame( { "cat": [1, 1, 1, 2, 2, 2, 2, 3, 3, 3], "max_col": [12, 24, 36, 15, 50, 50, 45, 20, 40, 60], "other_col": [25, 50, 75, 125, 150, 175, 200, 225, 250, 275], } ) # 按分组列升序、最大值列降序、顺序列升序排序,再取每个分组的第一条记录 result = df.sort( ["cat", "max_col", "other_col"], descending=[False, True, False] ).unique("cat", keep="first") print(result)
输出结果和你原代码一致:
shape: (3, 3) ┌─────┬─────────┬───────────┐ │ cat ┆ max_col ┆ other_col │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ i64 │ ╞═════╪═════════╪═══════════╡ │ 1 ┆ 36 ┆ 75 │ │ 2 ┆ 50 ┆ 150 │ │ 3 ┆ 60 ┆ 275 │ └─────┴─────────┴───────────┘
封装成通用函数
如果需要更灵活的复用(仅传入分组列和最大值列,可选指定顺序列),可以封装成函数:
def get_first_max_row(df, group_col, max_col, order_col=None): # 若未指定顺序列,用原数据的行号作为顺序依据(保证取原数据中首次出现的最大值) if order_col is None: df = df.with_row_index("row_order") order_col = "row_order" sorted_df = df.sort( [group_col, max_col, order_col], descending=[False, True, False] ).unique(group_col, keep="first") # 自动移除临时行号列 if order_col == "row_order": sorted_df = sorted_df.drop("row_order") return sorted_df # 使用示例:仅传入分组列和最大值列,自动按原数据顺序取首次出现的最大值 result = get_first_max_row(df, "cat", "max_col") # 也可以指定顺序列(如时间列) result = get_first_max_row(df, "cat", "max_col", "other_col")
方法优势
- 代码更简洁:无需嵌套两次
filter,逻辑直观 - 性能更优:排序+去重的操作在Polars中经过高度优化,比多次窗口过滤效率更高
- 灵活性强:可直接指定顺序列,或默认按原数据行号取首次出现的记录
内容的提问来源于stack exchange,提问作者niko86
相关产品推荐
相关产品推荐

