Polars链式表达式无法访问中间创建的index列问题
解决Polars链式表达式中Join子句无法访问新增index列的问题
错误原因
你链式写法里的df.select(...)调用的是原始未处理的DataFrame,它还没经过.with_row_index()添加index列,自然会抛出ColumnNotFoundError。
解决方法
方法一:用.pipe()传递已处理的DataFrame
通过pipe方法把前面添加了index列的DataFrame传递给Join的子查询,确保子查询使用的是已处理后的版本:
( df .with_row_index() .pipe(lambda indexed_df: indexed_df.join( indexed_df .select(["index", "event"]) .filter(pl.col("event") == 1) .with_row_index("event_id") .drop("event"), on="index", how="left" )) )
方法二:用窗口函数简化实现(无需Join)
不需要复杂的Join操作,直接用窗口函数就能生成目标event_id,代码更简洁高效:
( df .with_row_index() .with_columns( pl.when(pl.col("event") == 1) .then(pl.col("event").cum_count().over(pl.col("event") == 1) - 1) .alias("event_id") ) )
该逻辑中:
pl.col("event").cum_count().over(pl.col("event") == 1)会在event=1的分组内累计计数- 减1是让
event_id从0开始编号 event=0的行自动返回null,完全符合预期输出
验证结果
两种方法执行后都能得到你预期的输出:
shape: (4, 5) ┌───────┬───────┬─────┬─────┬──────────┐ │ index ┆ event ┆ foo ┆ boo ┆ event_id │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ u32 ┆ i64 ┆ i64 ┆ i64 ┆ u32 │ ╞═══════╪═══════╪═════╪═════╪══════════╡ │ 0 ┆ 0 ┆ 1 ┆ 2 ┆ null │ │ 1 ┆ 1 ┆ 2 ┆ 3 ┆ 0 │ │ 2 ┆ 1 ┆ 3 ┆ 4 ┆ 1 │ │ 3 ┆ 0 ┆ 4 ┆ 5 ┆ null │ └───────┴───────┴─────┴─────┴──────────┘
内容的提问来源于stack exchange,提问作者Kevin Li
相关产品推荐
相关产品推荐

