You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Polars中基于多列最大值条件获取对应x列值的最优方法

最优Polars实现方案:获取多列最大值对应的x值(适配批量场景)

先回顾下我们的原始数据:

import polars as pl

df = pl.DataFrame(dict(x=[0, 1, 2, 3], y=[5, 2, 3, 3],z=[4,7,8,2]))

对应的DataFrame输出:

shape: (4, 3)
┌─────┬─────┬─────┐
│ x   ┆ y   ┆ z   │
│ --- ┆ --- ┆ --- │
│ i64 ┆ i64 ┆ i64 │
╞═════╪═════╪═════╡
│ 0   ┆ 5   ┆ 4   │
│ 1   ┆ 2   ┆ 7   │
│ 2   ┆ 3   ┆ 8   │
│ 3   ┆ 3   ┆ 2   │
└─────┴─────┴─────┘

需求很明确:要批量找到**y、z(以及数百列)**取最大值时对应的x值,支持两种输出格式。下面是两个高效的Polars原生实现方案:

方案一:长表格式(shape: (n,2))

这种格式适合后续需要做分组分析的场景,完全自动适配任意多列,不需要手动指定列名:

result_long = (
    df
    # 把宽表转成长格式,x作为标识列,其他列转为variable-value对
    .melt(id_vars="x", value_vars=df.columns[1:])
    # 按列名分组
    .group_by("variable")
    .agg(
        # 过滤出当前列值为最大值的行,取对应的第一个x值
        pl.col("x").filter(pl.col("value") == pl.col("value").max()).first()
    )
    # 重命名列名符合需求
    .rename({"variable": "column", "x": "x_value"})
)

print(result_long)

输出结果:

shape: (2, 2)
┌────────┬─────────┐
│ column ┆ x_value │
│ ---    ┆ ---     │
│ str    ┆ i64     │
╞════════╪═════════╡
│ y      ┆ 0       │
│ z      ┆ 2       │
└────────┴─────────┘

方案二:宽表格式(shape: (1, n))

如果需要保持原列名作为结果的列名,用这个方案更直观,同样支持批量列处理:

result_wide = df.select(
    [
        # 对每个目标列,过滤出最大值对应的x值,用原列名作为别名
        pl.col("x").filter(pl.col(col) == pl.col(col).max()).first().alias(col)
        for col in df.columns[1:]
    ]
)

print(result_wide)

输出结果:

shape: (1, 2)
┌─────┬─────┐
│ y   ┆ z   │
│ --- ┆ --- │
│ i64 ┆ i64 │
╞═════╪═════╡
│ 0   ┆ 2   │
└─────┴─────┘

补充说明

  • 如果某列存在多个最大值对应的x值,可以把代码中的first()替换成list(),这样结果会返回所有符合条件的x值列表。
  • 两个方案都是基于Polars的向量化表达式实现,性能拉满,完全适配数百列的大规模数据场景,不需要写低效的循环。

内容的提问来源于stack exchange,提问作者pwb2103

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.04 18:55:14