You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Polars实现标记行对应区间内指定列的最大值计算?

如何用Polars实现指定规则的列计算?

给定如下Polars DataFrame:

import polars as pl

df = pl.DataFrame({
    "Column A": [2, 3, 1, 4, 1, 3, 3, 2, 1, 0],
    "Column B": [
        "Life", None, None, None, "Death", None, 
        "Life", None, None, "Death"
    ]
})

数据结构:

shape: (10, 2)
┌──────────┬──────────┐
│ Column A ┆ Column B │
│ ---      ┆ ---      │
│ i64      ┆ str      │
╞══════════╪══════════╡
│ 2        ┆ Life     │
│ 3        ┆ null     │
│ 1        ┆ null     │
│ 4        ┆ null     │
│ 1        ┆ Death    │
│ 3        ┆ null     │
│ 3        ┆ Life     │
│ 2        ┆ null     │
│ 1        ┆ null     │
│ 0        ┆ Death    │
└──────────┴──────────┘

需求:新增一列Column C,规则为:

  • 当Column B的值为'Life'时,Column C填入从该行到下一个Column B为'Death'的行之间Column A的最大值;
  • 其余行的Column C设为null。

预期结果:

shape: (10, 3)
┌──────────┬──────────┬──────────┐
│ Column A ┆ Column B ┆ Column C │
│ ---      ┆ ---      ┆ ---      │
│ i64      ┆ str      ┆ f64      │
╞══════════╪══════════╪══════════╡
│ 2        ┆ Life     ┆ 4.0      │
│ 3        ┆ null     ┆ null     │
│ 1        ┆ null     ┆ null     │
│ 4        ┆ null     ┆ null     │
│ 1        ┆ Death    ┆ null     │
│ 3        ┆ null     ┆ null     │
│ 3        ┆ Life     ┆ 3.0      │
│ 2        ┆ null     ┆ null     │
│ 1        ┆ null     ┆ null     │
│ 0        ┆ Death    ┆ null     │
└──────────┴──────────┴──────────┘

实现方案

可以通过标记分组区间 + 窗口函数计算最大值的方式实现,具体代码如下:

import polars as pl

df = pl.DataFrame({
    "Column A": [2, 3, 1, 4, 1, 3, 3, 2, 1, 0],
    "Column B": [
        "Life", None, None, None, "Death", None, 
        "Life", None, None, "Death"
    ]
})

result = df.with_columns(
    # 标记每个Life到Death的区间:遇到Life时生成新组ID,后续行继承该ID直到下一个Life
    group_id=pl.when(pl.col("Column B") == "Life")
               .then(pl.int_range(0, pl.count()))
               .forward_fill()
).with_columns(
    # 仅在Life行填充对应区间的Column A最大值,其余行设为null
    Column_C=pl.when(pl.col("Column B") == "Life")
               .then(pl.col("Column A").max().over("group_id"))
               .otherwise(None)
).drop("group_id")

print(result)

逻辑说明

  1. 标记分组区间:通过when(...).then(pl.int_range(...)).forward_fill()生成group_id,每遇到一个Life就分配一个递增的新ID,后续行自动继承该ID,直到下一个Life出现,以此划分出每个Life到下一个Life之间的区间(正好对应Life到下一个Death的范围)。
  2. 计算区间最大值:利用over("group_id")窗口函数,计算每个分组内Column A的最大值,再通过when条件仅在Column B为Life的行保留该值,其余行设为None。
  3. 清理临时列:最后删除用于分组的group_id临时列,得到目标结果。

运行代码后即可得到符合预期的DataFrame。

内容的提问来源于stack exchange,提问作者Levi Santos

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.29 10:33:16