You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中创建条件递增列?

Polars 实现条件递增列

要实现仅当code等于'L'时才递增的分组列(从1开始),可以通过以下步骤修改代码:

核心思路

  1. 先标记满足条件的行(code == 'L')为1,其余为0;
  2. 按file分组对标记列做累加,得到每个分组内截至当前行的'L'出现次数;
  3. 按需处理非条件行的数值(向前填充继承上一次的递增结果,或保持初始值1)。

完整代码

import polars as pl

df = pl.DataFrame({
    'file': ['a.txt','a.txt','a.txt','a.txt','b.txt','b.txt','c.txt','c.txt','c.txt','c.txt','c.txt'],
    'code': ['X','Y','Z','L','A','A','B','L','C','L','X']
})

# 生成条件递增列
df = df.with_columns(
    pl.when(pl.col('code') == 'L').then(1).otherwise(0)
      .cumsum()
      .over('file')
      .forward_fill()  # 非L行继承上一个L行的计数
      .fill_null(1)    # 无L的分组初始值设为1
      .alias('rrr')
)

print(df)

输出结果

shape: (11, 3)
┌─────────┬──────┬─────┐
│ file    ┆ code ┆ rrr │
│ ---     ┆ ---  ┆ --- │
│ str     ┆ str  ┆ i64 │
╞═════════╪══════╪═════╡
│ a.txt   ┆ X    ┆ 1   │
│ a.txt   ┆ Y    ┆ 1   │
│ a.txt   ┆ Z    ┆ 1   │
│ a.txt   ┆ L    ┆ 2   │
│ b.txt   ┆ A    ┆ 1   │
│ b.txt   ┆ A    ┆ 1   │
│ c.txt   ┆ B    ┆ 1   │
│ c.txt   ┆ L    ┆ 2   │
│ c.txt   ┆ C    ┆ 2   │
│ c.txt   ┆ L    ┆ 3   │
│ c.txt   ┆ X    ┆ 3   │
└─────────┴──────┴─────┘

可选调整

如果希望仅L行显示递增序号,其余行显示0,只需去掉.forward_fill().fill_null(1):

df = df.with_columns(
    pl.when(pl.col('code') == 'L').then(1).otherwise(0)
      .cumsum()
      .over('file')
      .alias('rrr')
)

对应的输出:

shape: (11, 3)
┌─────────┬──────┬─────┐
│ file    ┆ code ┆ rrr │
│ ---     ┆ ---  ┆ --- │
│ str     ┆ str  ┆ i64 │
╞═════════╪══════╪═════╡
│ a.txt   ┆ X    ┆ 0   │
│ a.txt   ┆ Y    ┆ 0   │
│ a.txt   ┆ Z    ┆ 0   │
│ a.txt   ┆ L    ┆ 1   │
│ b.txt   ┆ A    ┆ 0   │
│ b.txt   ┆ A    ┆ 0   │
│ c.txt   ┆ B    ┆ 0   │
│ c.txt   ┆ L    ┆ 1   │
│ c.txt   ┆ C    ┆ 1   │
│ c.txt   ┆ L    ┆ 2   │
│ c.txt   ┆ X    ┆ 2   │
└─────────┴──────┴─────┘

内容的提问来源于stack exchange,提问作者lmocsi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 03:50:01