You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何快速去除Polars DataFrame中整数列表列的连续重复项?

Polars去除列表列中连续重复项的高效方法

如果你的Polars DataFrame里有一列存的是长度可变的整数列表,要去掉每行列表里的连续重复项,优先用纯Polars内置表达式,这是性能最优的方案,完全避开Python UDF的开销,适合大数据量场景。

具体实现步骤

  1. 先构造示例DataFrame:
import polars as pl

df = pl.DataFrame({
    "nums": [
        [5, 5, 5, 4, 4, 5, 5, 6, 6, 7],
        [3, 4, 4, 5, 6],
        [2, 2],
        [1]
    ]
})
  1. 用Polars列表表达式处理:
df = df.with_columns(
    pl.col("nums").list.eval(
        pl.element().filter(pl.element().diff().is_not_null() | pl.element().first())
    ).alias("deduped_nums")
)

逻辑解释

  • pl.element().diff():计算列表中每个元素与前一个元素的差值,连续重复的元素差值为0,第一个元素的差值为null
  • pl.element().diff().is_not_null():筛选出差值不为null且不为0的元素(也就是和前一个元素不同的项)
  • | pl.element().first():把列表的第一个元素强制保留,避免被过滤掉
  • 最终过滤后的列表就是去除了连续重复项的结果

输出结果

执行后打印DataFrame会得到:

shape: (4, 2)
┌───────────────────────────────────┬───────────────────┐
│ nums                              ┆ deduped_nums      │
│ ---                               ┆ ---               │
│ list[i64]                         ┆ list[i64]         │
╞═══════════════════════════════════╪═══════════════════╡
│ [5, 5, 5, 4, 4, 5, 5, 6, 6, 7]   ┆ [5, 4, 5, 6, 7]   │
│ [3, 4, 4, 5, 6]                   ┆ [3, 4, 5, 6]      │
│ [2, 2]                            ┆ [2]               │
│ [1]                               ┆ [1]               │
└───────────────────────────────────┴───────────────────┘

不推荐的方案(小数据量可用)

如果是小数据集,也可以用Python函数配合map_elements实现,但大数据量下性能会差很多,因为要频繁在Polars引擎和Python上下文之间切换:

def dedupe_consecutive(lst):
    if not lst:
        return lst
    res = [lst[0]]
    for num in lst[1:]:
        if num != res[-1]:
            res.append(num)
    return res

df = df.with_columns(
    pl.col("nums").map_elements(dedupe_consecutive).alias("deduped_nums")
)

内容的提问来源于stack exchange,提问作者Hendrik Wiese

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.22 19:42:16