You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中按行执行自定义累积计算?

在Polars中实现自定义递推累积计算

你需要用自定义函数f(prev, curr) = prev * 2 + curr对Polars DataFrame的some_col做递推累积计算,替换原列值得到预期结果。Polars提供了cumulative_eval方法,能优雅实现这种依赖前序结果的计算,比手动循环更符合Polars范式,效率也更高。

完整实现代码

import polars as pl

def f(prev, curr):
    return prev * 2 + curr

# 构造示例DataFrame
df = pl.DataFrame({
    "some_col": [7, 3, 9, 2],
    "other_col": ["...", "", "", ""]
})

# 执行递推累积计算
df = df.with_columns(
    pl.col("some_col").cumulative_eval(
        lambda acc, curr: f(acc, curr),
        initial=pl.col("some_col").first()
    ).alias("some_col")
)

print(df)

代码说明

  • cumulative_eval是Polars专门处理递推累积的方法:
    • 第一个参数是处理逻辑的匿名函数,acc代表上一步的累积结果,curr代表当前行的列值,直接调用自定义函数f即可完成计算。
    • initial指定累积的起始值,这里取some_col的第一行数据作为初始值。
  • 用alias指定列名,这里直接替换原some_col;如果需要保留原列,可以改成新列名(比如result_col)。

输出结果

运行后得到的DataFrame完全符合预期:

shape: (4, 2)
┌──────────┬────────────┐
│ some_col ┆ other_col  │
│ ---      ┆ ---        │
│ i64      ┆ str        │
╞══════════╪════════════╡
│ 7        ┆ ...        │
│ 17       ┆            │
│ 43       ┆            │
│ 88       ┆            │
└──────────┴────────────┘

扩展说明

不管你的自定义函数逻辑多复杂,只要它能接收prev(前序累积值)和curr(当前值)两个参数并返回计算结果,都可以直接套用到这个模式里。而且cumulative_eval底层基于Polars的优化机制,数据量越大,相比手动循环的性能优势越明显。

内容的提问来源于stack exchange,提问作者ldmat

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 22:40:00