如何在Polars中应用带多个参数的自定义函数实现行级计算
Polars 实现对应计算的方法
你可以根据场景选择以下两种实现方式,最终计算结果和Pandas代码完全一致:
方案1:原生向量化表达式(强烈推荐)
Polars 原生向量化运算运行在Rust层,没有Python逐行调用的性能开销,是最优实现方式。你的示例逻辑不需要自定义函数,直接写表达式即可:
import polars as pl # 构造测试数据 df = pl.DataFrame({ "a":[1,2,3,4,5], "b":[2,3,4,5,6], "c":[3,4,5,6,7] }) shift_len = 3 df = df.with_columns( d = pl.col("a") + pl.col("b") * shift_len, e = pl.col("b") - shift_len )
运行后得到的d、e列和Pandas处理结果完全匹配:第一行d=7、e=-1,第二行d=11、e=0,逐行对应。
方案2:复用现有Python自定义函数
如果你的自定义函数逻辑复杂,无法直接拆成Polars表达式,可以用逐行映射的方式实现,效果等同于Pandas的apply(axis=1, result_type="expand"):
def fun(a,b,shift_len): return a+b*shift_len,b-shift_len shift_len = 3 df = df.with_columns( pl.struct(["a", "b"]) .map_rows(lambda row: fun(row["a"], row["b"], shift_len)) .struct.rename_fields(["d", "e"]) .alias("tmp") ).unnest("tmp")
实现逻辑说明:
- 用
pl.struct把计算需要用到的a、b列打包为结构体传入逐行映射 map_rows逐行调用你的自定义函数,返回多值结果- 给返回的结构体字段重命名为目标列名d、e,最后用
unnest把结构体展开为独立列即可。
注意:逐行调用Python函数的方式性能远低于原生向量化实现,非必要不使用。
内容的提问来源于stack exchange,提问作者user3105812
相关产品推荐
相关产品推荐

