如何快速去除Polars DataFrame中整数列表列的连续重复项?
Polars去除列表列中连续重复项的高效方法
如果你的Polars DataFrame里有一列存的是长度可变的整数列表,要去掉每行列表里的连续重复项,优先用纯Polars内置表达式,这是性能最优的方案,完全避开Python UDF的开销,适合大数据量场景。
具体实现步骤
- 先构造示例DataFrame:
import polars as pl df = pl.DataFrame({ "nums": [ [5, 5, 5, 4, 4, 5, 5, 6, 6, 7], [3, 4, 4, 5, 6], [2, 2], [1] ] })
- 用Polars列表表达式处理:
df = df.with_columns( pl.col("nums").list.eval( pl.element().filter(pl.element().diff().is_not_null() | pl.element().first()) ).alias("deduped_nums") )
逻辑解释
pl.element().diff():计算列表中每个元素与前一个元素的差值,连续重复的元素差值为0,第一个元素的差值为nullpl.element().diff().is_not_null():筛选出差值不为null且不为0的元素(也就是和前一个元素不同的项)| pl.element().first():把列表的第一个元素强制保留,避免被过滤掉- 最终过滤后的列表就是去除了连续重复项的结果
输出结果
执行后打印DataFrame会得到:
shape: (4, 2) ┌───────────────────────────────────┬───────────────────┐ │ nums ┆ deduped_nums │ │ --- ┆ --- │ │ list[i64] ┆ list[i64] │ ╞═══════════════════════════════════╪═══════════════════╡ │ [5, 5, 5, 4, 4, 5, 5, 6, 6, 7] ┆ [5, 4, 5, 6, 7] │ │ [3, 4, 4, 5, 6] ┆ [3, 4, 5, 6] │ │ [2, 2] ┆ [2] │ │ [1] ┆ [1] │ └───────────────────────────────────┴───────────────────┘
不推荐的方案(小数据量可用)
如果是小数据集,也可以用Python函数配合map_elements实现,但大数据量下性能会差很多,因为要频繁在Polars引擎和Python上下文之间切换:
def dedupe_consecutive(lst): if not lst: return lst res = [lst[0]] for num in lst[1:]: if num != res[-1]: res.append(num) return res df = df.with_columns( pl.col("nums").map_elements(dedupe_consecutive).alias("deduped_nums") )
内容的提问来源于stack exchange,提问作者Hendrik Wiese
相关产品推荐
相关产品推荐

