如何在Python Polars中实现依赖同列前值的递归计算?
解决Polars中依赖前序结果的递推列生成问题
要实现你描述的递推逻辑(第一行返回"A",后续行继承上一行的Result值),核心是利用Polars的cumulative_eval函数——它专门用于处理需要逐行依赖前序计算结果的场景,完美解决with_columns无法引用未创建列的问题。
基础示例代码
假设你的DataFrame结构如下:
import polars as pl df = pl.DataFrame({ "Index": [0, 1, 2, 3, 4] })
生成Result列的代码:
df = df.with_columns( # 定义递推规则:索引为0时取"A",否则取上一行的Result pl.when(pl.int_range(0, pl.count()) == 0) .then(pl.lit("A")) .otherwise(pl.col("Result").shift()) # 逐行执行递推计算 .cumulative_eval() .alias("Result") )
执行后输出的DataFrame:
shape: (5, 2) ┌───────┬────────┐ │ Index ┆ Result │ │ --- ┆ --- │ │ i64 ┆ str │ ╞═══════╪════════╡ │ 0 ┆ A │ │ 1 ┆ A │ │ 2 ┆ A │ │ 3 ┆ A │ │ 4 ┆ A │ └───────┴────────┘
扩展到复杂递推场景
如果你的实际需求是更复杂的递推(比如Result[i] = Result[i-1] + str(i)),只需修改otherwise里的逻辑即可:
df = df.with_columns( pl.when(pl.int_range(0, pl.count()) == 0) .then(pl.lit("A")) .otherwise(pl.col("Result").shift() + pl.col("Index").cast(pl.Utf8)) .cumulative_eval() .alias("Result") )
输出结果:
shape: (5, 2) ┌───────┬────────┐ │ Index ┆ Result │ │ --- ┆ --- │ │ i64 ┆ str │ ╞═══════╪════════╡ │ 0 ┆ A │ │ 1 ┆ A1 │ │ 2 ┆ A12 │ │ 3 ┆ A123 │ │ 4 ┆ A1234 │ └───────┴────────┘
为什么之前的方法无效?
你尝试的shift()、rolling()等方法都是基于原始列的静态数据计算,无法引用正在创建的列的动态结果。而cumulative_eval会逐行执行计算,每一步都可以复用前一行的计算结果,完全适配递推逻辑的需求。
内容的提问来源于stack exchange,提问作者Danilo Setton
相关产品推荐
相关产品推荐

