优化Polars中布尔列表过滤列表列且保留原长的性能方案
Polars 高性能实现布尔列表过滤字符串列表(保留原长度)
目标输出
我们需要生成如下结构的DataFrame,新增的filtered_strings列会在identity_vector对应布尔值为False的位置填充null,同时保留原列表长度:
shape: (2, 3) ┌──────────────────┬──────────────────┬───────────────────┐ │ identity_vector ┆ string_vector ┆ filtered_strings │ │ --- ┆ --- ┆ --- │ │ list[bool] ┆ list[str] ┆ list[str | null] │ ╞══════════════════╪══════════════════╪═══════════════════╡ │ [true, false] ┆ ["name1", "name2"] ┆ ["name1", null] │ │ [false, true] ┆ ["name3", "name4"] ┆ [null, "name4"] │ └──────────────────┴──────────────────┴───────────────────┘
高性能实现方案
避免使用性能低下的map_elements(逐行Python循环),改用Polars原生的向量化数组操作,代码如下:
import polars as pl df = pl.DataFrame({ 'identity_vector': [[True, False], [False, True]], 'string_vector': [['name1', 'name2'], ['name3', 'name4']] }) # 生成filtered_strings列 df = df.with_columns( filtered_strings=pl.arr.zip("identity_vector", "string_vector") .arr.eval(pl.element().tuple().get(1).when(pl.element().tuple().get(0)).otherwise(None)) ) print(df)
方案说明
pl.arr.zip:将identity_vector和string_vector按位置打包成元组列表,例如第一行变为[(True, 'name1'), (False, 'name2')]arr.eval:对每个元组执行向量化判断——当元组第一个元素(布尔值)为True时保留第二个元素(字符串),否则填充null- 整个过程完全在Polars查询引擎内执行,无Python循环开销,大数据量下性能远优于
map_elements
内容的提问来源于stack exchange,提问作者yz_jc
相关产品推荐
相关产品推荐

