You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中根据分组首行的rt条件更新对应字段值

如何在Polars中根据分组首行的rt条件更新对应字段值

这个需求用Polars的窗口函数加条件判断就能轻松实现,我来给你一步步拆解怎么做:

完整实现代码

import polars as pl

df = pl.DataFrame({
    'file':['a','a','a','a','b','b'],
    'ru':['fe','fe','ev','ev','ba','br'],
    'rt':[0,0,1,1,1,0],
})

# 实现核心逻辑
result = df.with_columns(
    # 为每个分组获取首行的ru和rt值,作为临时参考字段
    first_ru=pl.col('ru').first().over('file'),
    first_rt=pl.col('rt').first().over('file')
).with_columns(
    # 条件替换:如果分组首行rt是0,就用首行值替换整组的ru和rt;否则保留原数据
    ru=pl.when(pl.col('first_rt') == 0)
         .then(pl.col('first_ru'))
         .otherwise(pl.col('ru')),
    rt=pl.when(pl.col('first_rt') == 0)
         .then(pl.col('first_rt'))
         .otherwise(pl.col('rt'))
).drop(['first_ru', 'first_rt'])  # 清理临时字段

print(result)

代码逻辑解释

  1. 获取分组首行参考值:通过over('file')窗口函数,为每一行生成first_ru和first_rt两个临时字段,分别存储当前file分组的第一行ru和rt值。这样每一行都能快速拿到自己分组的首行基准值。
  2. 条件替换字段:用pl.when().then().otherwise()做分支判断:
    • 当分组首行的rt值(也就是first_rt)为0时,把当前行的ru替换成first_ru,rt替换成first_rt;
    • 如果首行rt不为0,就保持原字段的值不变。
  3. 清理临时字段:最后删掉我们中间用的first_ru和first_rt,得到干净的结果。

运行结果

执行后你会得到完全符合预期的输出:

shape: (6, 3)
┌──────┬─────┬─────┐
│ file ┆ ru  ┆ rt  │
│ ---  ┆ --- ┆ --- │
│ str  ┆ str ┆ i64 │
╞══════╪═════╪═════╡
│ a    ┆ fe  ┆ 0   │
│ a    ┆ fe  ┆ 0   │
│ a    ┆ fe  ┆ 0   │
│ a    ┆ fe  ┆ 0   │
│ b    ┆ ba  ┆ 1   │
│ b    ┆ br  ┆ 0   │
└──────┴─────┴─────┘

备注:内容来源于stack exchange,提问作者lmocsi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.17 09:02:59