You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars链式表达式无法访问中间创建的index列问题

解决Polars链式表达式中Join子句无法访问新增index列的问题

错误原因

你链式写法里的df.select(...)调用的是原始未处理的DataFrame,它还没经过.with_row_index()添加index列,自然会抛出ColumnNotFoundError。

解决方法

方法一:用.pipe()传递已处理的DataFrame

通过pipe方法把前面添加了index列的DataFrame传递给Join的子查询,确保子查询使用的是已处理后的版本:

(
    df
    .with_row_index()
    .pipe(lambda indexed_df: indexed_df.join(
        indexed_df
        .select(["index", "event"])
        .filter(pl.col("event") == 1)
        .with_row_index("event_id")
        .drop("event"),
        on="index",
        how="left"
    ))
)

方法二:用窗口函数简化实现(无需Join)

不需要复杂的Join操作,直接用窗口函数就能生成目标event_id,代码更简洁高效:

(
    df
    .with_row_index()
    .with_columns(
        pl.when(pl.col("event") == 1)
        .then(pl.col("event").cum_count().over(pl.col("event") == 1) - 1)
        .alias("event_id")
    )
)

该逻辑中:

  • pl.col("event").cum_count().over(pl.col("event") == 1)会在event=1的分组内累计计数
  • 减1是让event_id从0开始编号
  • event=0的行自动返回null,完全符合预期输出

验证结果

两种方法执行后都能得到你预期的输出:

shape: (4, 5)
┌───────┬───────┬─────┬─────┬──────────┐
│ index ┆ event ┆ foo ┆ boo ┆ event_id │
│ ---   ┆ ---   ┆ --- ┆ --- ┆ ---      │
│ u32   ┆ i64   ┆ i64 ┆ i64 ┆ u32      │
╞═══════╪═══════╪═════╪═════╪══════════╡
│ 0     ┆ 0     ┆ 1   ┆ 2   ┆ null     │
│ 1     ┆ 1     ┆ 2   ┆ 3   ┆ 0        │
│ 2     ┆ 1     ┆ 3   ┆ 4   ┆ 1        │
│ 3     ┆ 0     ┆ 4   ┆ 5   ┆ null     │
└───────┴───────┴─────┴─────┴──────────┘

内容的提问来源于stack exchange,提问作者Kevin Li

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 18:26:06