如何用Polars表达式对外部列表/NumPy数组进行索引与切片?
用Polars表达式替代map_elements处理外部NumPy数组切片
示例数据与初始实现
先定义测试用的外部NumPy数组列表和Polars DataFrame:
import polars as pl import numpy as np # 外部NumPy数组列表 audio_arrays = [ np.array([10, 20, 30, 40, 50]), np.array([100, 200, 300, 400]), np.array([1000, 2000]) ] # 示例Polars DataFrame df = pl.DataFrame({ "audio_idx": [0, 0, 1, 2], "i_start": [0, 2, 1, 0], "i_stop": [2, 5, 3, 2] })
原来基于map_elements的实现(逐行处理,性能较低):
# 初始map_elements实现 result_df = df.with_columns( pl.struct(["audio_idx", "i_start", "i_stop"]) .map_elements(lambda x: audio_arrays[x["audio_idx"]][x["i_start"]:x["i_stop"]], return_dtype=pl.List(pl.Int64)) .alias("sliced_audio") )
期望输出结果:
shape: (4, 4) ┌──────────┬─────────┬────────┬──────────────────┐ │ audio_idx ┆ i_start ┆ i_stop ┆ sliced_audio │ │ --- ┆ --- ┆ --- ┆ --- │ │ i64 ┆ i64 ┆ i64 ┆ list[i64] │ ╞══════════╪═════════╪════════╪══════════════════╡ │ 0 ┆ 0 ┆ 2 ┆ [10, 20] │ │ 0 ┆ 2 ┆ 5 ┆ [30, 40, 50] │ │ 1 ┆ 1 ┆ 3 ┆ [200, 300] │ │ 2 ┆ 0 ┆ 2 ┆ [1000, 2000] │ └──────────┴─────────┴────────┴──────────────────┘
Polars表达式优化实现
使用pl.map_batches替代map_elements,按批次处理数据,同时保留外部数组列表以节省内存:
# 基于Polars表达式的优化实现 result_df = df.with_columns( pl.map_batches( ["audio_idx", "i_start", "i_stop"], lambda idx_col, start_col, stop_col: pl.Series( [audio_arrays[i][s:e] for i, s, e in zip(idx_col.to_list(), start_col.to_list(), stop_col.to_list())] ), return_dtype=pl.List(pl.Int64) ).alias("sliced_audio") )
方案优势
- 性能更优:
map_batches按批次处理数据,避免了map_elements的逐行遍历开销 - 内存高效:直接引用外部的
audio_arrays列表,无需将大型NumPy数组导入Polars内部存储 - 符合表达式API要求:完全基于Polars的表达式体系实现,没有依赖逐行的lambda处理
内容的提问来源于stack exchange,提问作者Brian Provost
相关产品推荐
相关产品推荐

