如何用Polars实现标记行对应区间内指定列的最大值计算?
如何用Polars实现指定规则的列计算?
给定如下Polars DataFrame:
import polars as pl df = pl.DataFrame({ "Column A": [2, 3, 1, 4, 1, 3, 3, 2, 1, 0], "Column B": [ "Life", None, None, None, "Death", None, "Life", None, None, "Death" ] })
数据结构:
shape: (10, 2) ┌──────────┬──────────┐ │ Column A ┆ Column B │ │ --- ┆ --- │ │ i64 ┆ str │ ╞══════════╪══════════╡ │ 2 ┆ Life │ │ 3 ┆ null │ │ 1 ┆ null │ │ 4 ┆ null │ │ 1 ┆ Death │ │ 3 ┆ null │ │ 3 ┆ Life │ │ 2 ┆ null │ │ 1 ┆ null │ │ 0 ┆ Death │ └──────────┴──────────┘
需求:新增一列Column C,规则为:
- 当
Column B的值为'Life'时,Column C填入从该行到下一个Column B为'Death'的行之间Column A的最大值; - 其余行的
Column C设为null。
预期结果:
shape: (10, 3) ┌──────────┬──────────┬──────────┐ │ Column A ┆ Column B ┆ Column C │ │ --- ┆ --- ┆ --- │ │ i64 ┆ str ┆ f64 │ ╞══════════╪══════════╪══════════╡ │ 2 ┆ Life ┆ 4.0 │ │ 3 ┆ null ┆ null │ │ 1 ┆ null ┆ null │ │ 4 ┆ null ┆ null │ │ 1 ┆ Death ┆ null │ │ 3 ┆ null ┆ null │ │ 3 ┆ Life ┆ 3.0 │ │ 2 ┆ null ┆ null │ │ 1 ┆ null ┆ null │ │ 0 ┆ Death ┆ null │ └──────────┴──────────┴──────────┘
实现方案
可以通过标记分组区间 + 窗口函数计算最大值的方式实现,具体代码如下:
import polars as pl df = pl.DataFrame({ "Column A": [2, 3, 1, 4, 1, 3, 3, 2, 1, 0], "Column B": [ "Life", None, None, None, "Death", None, "Life", None, None, "Death" ] }) result = df.with_columns( # 标记每个Life到Death的区间:遇到Life时生成新组ID,后续行继承该ID直到下一个Life group_id=pl.when(pl.col("Column B") == "Life") .then(pl.int_range(0, pl.count())) .forward_fill() ).with_columns( # 仅在Life行填充对应区间的Column A最大值,其余行设为null Column_C=pl.when(pl.col("Column B") == "Life") .then(pl.col("Column A").max().over("group_id")) .otherwise(None) ).drop("group_id") print(result)
逻辑说明
- 标记分组区间:通过
when(...).then(pl.int_range(...)).forward_fill()生成group_id,每遇到一个Life就分配一个递增的新ID,后续行自动继承该ID,直到下一个Life出现,以此划分出每个Life到下一个Life之间的区间(正好对应Life到下一个Death的范围)。 - 计算区间最大值:利用
over("group_id")窗口函数,计算每个分组内Column A的最大值,再通过when条件仅在Column B为Life的行保留该值,其余行设为None。 - 清理临时列:最后删除用于分组的
group_id临时列,得到目标结果。
运行代码后即可得到符合预期的DataFrame。
内容的提问来源于stack exchange,提问作者Levi Santos
相关产品推荐
相关产品推荐

