You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars:如何并行化仅用Polars表达式的Lambda?代码为何单核心运行?

问题原因与解决方案

为什么代码只在单核心运行?

你的代码里用了map_elements(lambda x: x.cast(pl.String)),map_elements本质是逐行执行的用户自定义函数(UDF),Polars无法对这类操作做并行优化,只能在单个核心上运行,这就是导致单核心运行的核心原因。

优化后的多核并行代码

要实现将列表转成目标字符串的需求,完全可以用Polars原生的向量化表达式,不需要依赖map_elements,这样就能充分利用多核性能:

import polars as pl

df = pl.DataFrame(dict(ent=['a', 'b'], doc_ids=[[2,3], [3]]))
df = (df.lazy()
    .with_columns(
        pl.concat_str(
            pl.lit('['),
            pl.col('doc_ids').list.eval(pl.element().cast(pl.String)).list.join(', '),
            pl.lit(']')
        ).alias('docs_str')
    )
    .drop('doc_ids')
).collect()

代码说明

  • 用list.eval(pl.element().cast(pl.String))替代map_elements:list.eval是Polars原生的列表元素批量处理方法,支持向量化并行,能利用多核资源。
  • 后续的list.join(', ')和concat_str都是Polars原生的向量化操作,本身就支持并行执行。

更简洁的写法(Polars >=0.19.0)

如果你的Polars版本足够新,还可以用list.to_string()方法直接实现需求,代码更简洁:

import polars as pl

df = pl.DataFrame(dict(ent=['a', 'b'], doc_ids=[[2,3], [3]]))
df = (df.lazy()
    .with_columns(
        pl.col('doc_ids').list.to_string().alias('docs_str')
    )
    .drop('doc_ids')
).collect()

这个方法会直接把列表转成"[2, 3]"格式的字符串,完全符合需求,而且是纯原生向量化操作,并行效率拉满。

内容的提问来源于stack exchange,提问作者Tim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 08:15:44