如何高效用Polars生成每行含1...n列表的pl.Series?
如何高效创建每行包含1到n的列表的Polars Series?
我尝试用pl.int_range创建一个Polars Series,其中每行元素是包含1到对应a列值的列表,但运行以下代码时报错:
import polars as pl df = pl.DataFrame({ 'a': [1, 2, 3, 4, 5] }).select( pl.int_range(1, pl.col('a')+1) )
报错信息:
ComputeError: `end` must contain exactly one value, got 5 values
我可以用map_elements结合Python的range绕开这个问题,但担心性能不够理想:
df = pl.DataFrame({ 'a': [1, 2, 3, 4, 5] }) print(df.with_columns( pl.col('a').map_elements(lambda x: list(range(1, x+1))).alias('c') ))
运行结果:
shape: (5, 2) ┌─────┬───────────────┐ │ a ┆ c │ │ --- ┆ --- │ │ i64 ┆ list[i64] │ ╞═════╪═══════════════╡ │ 1 ┆ [1] │ │ 2 ┆ [1, 2] │ │ 3 ┆ [1, 2, 3] │ │ 4 ┆ [1, 2, ... 4] │ │ 5 ┆ [1, 2, ... 5] │ └─────┴───────────────┘
高效实现方式
使用Polars原生的pl.list.range函数(Polars 0.19.0及以上版本支持),这是完全矢量化的操作,性能远优于map_elements:
import polars as pl df = pl.DataFrame({'a': [1, 2, 3, 4, 5]}) result = df.with_columns( pl.list.range(1, pl.col('a') + 1).alias('c') ) print(result)
运行结果与之前一致,但避免了Python层面的循环开销,处理大数据量时性能提升明显。
如果使用的是较旧版本的Polars,可改用map_batches结合pl.int_range优化,效率也比map_elements更高:
import polars as pl df = pl.DataFrame({'a': [1, 2, 3, 4, 5]}) result = df.with_columns( pl.col('a').map_batches( lambda s: s.apply(lambda x: pl.int_range(1, x+1).to_list()), return_dtype=pl.List(pl.Int64) ).alias('c') ) print(result)
报错原因说明
pl.int_range的end参数要求是单个标量值,无法接收包含多个元素的列;而pl.list.range专门用于生成每行对应长度的列表,支持列作为参数,完美适配需求。
内容的提问来源于stack exchange,提问作者NedDasty
相关产品推荐
相关产品推荐

