如何在Polars延迟数据框中保留其他列并逐行执行自定义函数?
解决Polars Lazy DataFrame中保留额外列并逐行处理自定义函数的问题
要在Polars Lazy DataFrame中保留id这类非目标列,同时对指定列(如a/b/c)逐行运行自定义函数(比如softmax),核心是避免只针对单个列调用map_batches,而是将目标列打包成struct后批量处理,同时保留需要的列。以下是几种可行方案:
准备示例代码
先构造测试数据和softmax函数:
import polars as pl import numpy as np # 示例Lazy DataFrame df = pl.LazyFrame({ "id": [1, 2, 3], "a": [1.0, 2.0, 3.0], "b": [4.0, 5.0, 6.0], "c": [7.0, 8.0, 9.0] }) # 自定义softmax函数 def softmax(x): exp_x = np.exp(x - np.max(x, axis=1, keepdims=True)) return exp_x / np.sum(exp_x, axis=1, keepdims=True)
方案1:Struct打包+Map Batches(推荐,高性能)
将目标列打包成struct,通过map_batches批量处理,同时保留id列,最后展开结果:
result = df.select( pl.col("id"), pl.struct(["a", "b", "c"]).map_batches( lambda struct_col: pl.DataFrame( softmax(struct_col.to_numpy()), columns=["a_softmax", "b_softmax", "c_softmax"] ) ).alias("softmax_results") ).unnest("softmax_results") # 执行并查看结果 print(result.collect())
如果需要保留所有原列(包括a/b/c),改用with_columns即可:
result = df.with_columns( pl.struct(["a", "b", "c"]).map_batches( lambda struct_col: pl.DataFrame( softmax(struct_col.to_numpy()), columns=["a_softmax", "b_softmax", "c_softmax"] ) ).alias("softmax_results") ).unnest("softmax_results")
方案2:Map Rows(适合小数据集)
如果数据量不大,map_rows可以直接逐行处理,代码更直观,但性能不如批量处理:
result = df.map_rows( lambda row: { "id": row["id"], "a": softmax(np.array([row["a"], row["b"], row["c"]]).reshape(1, -1))[0][0], "b": softmax(np.array([row["a"], row["b"], row["c"]]).reshape(1, -1))[0][1], "c": softmax(np.array([row["a"], row["b"], row["c"]]).reshape(1, -1))[0][2] } ) print(result.collect())
方案3:替换原列(可选)
如果希望用softmax结果直接替换原列a/b/c,可以调整列名后替换:
result = df.with_columns( pl.struct(["a", "b", "c"]).map_batches( lambda struct_col: pl.DataFrame( softmax(struct_col.to_numpy()), columns=["a", "b", "c"] ) ).alias("temp") ).drop(["a", "b", "c"]).unnest("temp")
内容的提问来源于stack exchange,提问作者velochy
相关产品推荐
相关产品推荐

