如何使用Polars原生API过滤列表类型列中的指定元素?
问题描述
假设我有一个包含字符串列表类型列的Polars DataFrame:
┌─────────────────────────────────────────────────┐ │ words │ │ --- │ │ list[str] │ ╞═════════════════════════════════════════════════╡ │ ["i", "like", "the", "pizza"] │ │ ["the", "dog", "is", "runnig"] │ │ ["me", "and", "my", "friend", "are", "playing"] │ └─────────────────────────────────────────────────┘
我需要过滤每个列表中的停用词,目前用map_elements结合自定义lambda可以实现,但文档说明自定义UDF运行速度较慢,想找Polars原生API的解决方案。
原生API解决方案
Polars提供了原生的列表过滤方法,完全不需要依赖自定义UDF,性能更优。
方法1:使用list.filter(推荐,Polars 0.18.0+)
直接通过list.filter结合is_in的取反逻辑过滤停用词,这是最简洁高效的方式:
import polars as pl pl.Config(fmt_table_cell_list_len=8, fmt_str_lengths=80) df = pl.DataFrame({ "words": [["i", "like", "the", "pizza"], ["the", "dog", "is", "runnig"], ["me", "and", "my", "friend", "are", "playing"]] }) STOP_WORDS = ["the"] filtered_df = df.with_columns( pl.col("words").list.filter(~pl.element().is_in(STOP_WORDS)).alias("words") ) print(filtered_df)
执行后得到的结果与原方法一致:
┌─────────────────────────────────────────────────┐ │ words │ │ --- │ │ list[str] │ ╞═════════════════════════════════════════════════╡ │ ["i", "like", "pizza"] │ │ ["dog", "is", "runnig"] │ │ ["me", "and", "my", "friend", "are", "playing"] │ └─────────────────────────────────────────────────┘
方法2:使用list.eval(兼容旧版本Polars)
如果你的Polars版本低于0.18.0,list.filter尚未推出,可以用list.eval结合元素过滤表达式实现:
filtered_df = df.with_columns( pl.col("words").list.eval(pl.element().filter(~pl.element().is_in(STOP_WORDS))).alias("words") )
性能说明
原生API基于Polars的矢量化引擎实现,避免了Python层面的循环开销,在处理大数据集时,速度比map_elements快数倍甚至数十倍。
内容的提问来源于stack exchange,提问作者barak1412
相关产品推荐
相关产品推荐

