如何在Polars DataFrame中基于两列比较快速创建新字符串列?
Polars中通过比较两列快速生成含"+"/"-"的字符串列
最优解法:使用Polars原生表达式(推荐)
不需要用map_batches这类用户自定义函数,直接用Polars内置的when/then/otherwise矢量化操作,既简洁高效,还能避免类型推断错误:
import polars as pl df = pl.from_repr(""" ┌─────┬─────┐ │ a ┆ b │ │ --- ┆ --- │ │ i64 ┆ i64 │ ╞═════╪═════╡ │ 2 ┆ 20 │ │ 30 ┆ 3 │ └─────┴─────┘ """) # 生成目标列 result_df = df.with_columns( pl.when(pl.col("a") < pl.col("b")) .then("+") .otherwise("-") .alias("strand") ) print(result_df)
执行后直接得到目标结果:
┌─────┬─────┬────────┐ │ a ┆ b ┆ strand │ │ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ str │ ╞═════╪═════╪════════╡ │ 2 ┆ 20 ┆ + │ │ 30 ┆ 3 ┆ - │ └─────┴─────┴────────┘
若坚持使用map_batches的解决方案
如果一定要用map_batches,需要手动指定返回类型return_dtype=pl.String,让Polars明确输出类型:
result_df = df.with_columns( pl.map_batches( ["a", "b"], lambda s: "+" if s[0] < s[1] else "-", return_dtype=pl.String ).alias("strand") )
为什么原生表达式更好
- 矢量化操作比用户自定义函数性能高得多,处理大数据集时差距明显
- 无需手动指定类型,Polars能自动推断返回的字符串类型
- 代码可读性更强,符合Polars的惯用写法
内容的提问来源于stack exchange,提问作者darked89
相关产品推荐
相关产品推荐

