You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars中如何移除重复id对应的首行(首行值为其余行值之和)

Polars高效移除分组总和行的实现方法

原始DataFrame

import polars as pl

df = pl.DataFrame({
    "id": [1,2,2,2,2,3,3,3],
    "value": [5,6,1,2,3,30,10,20]
})

输出结果:

┌─────┬───────┐
│ id  ┆ value │
│ --- ┆ ---   │
│ i64 ┆ i64   │
╞═════╪═══════╡
│ 1   ┆ 5     │
│ 2   ┆ 6     │
│ 2   ┆ 1     │
│ 2   ┆ 2     │
│ 2   ┆ 3     │
│ 3   ┆ 30    │
│ 3   ┆ 10    │
│ 3   ┆ 20    │
└─────┴───────┘

需求说明

当同一id对应多行数据时,首行的value是其余行value的总和,需要移除这些作为总和的行,保留id仅一行的数据以及分组内的非总和行。

高效实现代码

利用Polars的窗口函数实现向量化处理,无需额外分组合并操作,性能更优:

result = df.with_row_index() \
    .filter(
        ~(
            # 判断是否是分组首行
            (pl.col("index") == pl.col("index").min().over("id")) &
            # 判断分组行数大于1(即存在需要移除的总和行)
            (pl.col("id").count().over("id") > 1) &
            # 判断当前value等于其余行的总和(分组总和减去自身)
            (pl.col("value") == pl.col("value").sum().over("id") - pl.col("value"))
        )
    ) \
    .drop("index")

print(result)

输出结果

┌─────┬───────┐
│ id  ┆ value │
│ --- ┆ ---   │
│ i64 ┆ i64   │
╞═════╪═══════╡
│ 1   ┆ 5     │
│ 2   ┆ 1     │
│ 2   ┆ 2     │
│ 2   ┆ 3     │
│ 3   ┆ 10    │
│ 3   ┆ 20    │
└─────┴───────┘

逻辑说明

  1. with_row_index():添加行索引,用于定位每个分组的首行(原始数据中总和行固定为分组首行);
  2. 窗口函数over("id"):针对每个id分组计算聚合值(首行索引、分组行数、分组总和);
  3. 过滤条件:排除同时满足「分组首行」「分组行数>1」「当前value等于其余行总和」的行;
  4. drop("index"):删除辅助索引列,得到最终结果。

这种方法全程基于Polars的向量化操作,处理大数据量时效率远高于循环或分组后二次合并的方式。

内容的提问来源于stack exchange,提问作者Julian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.09 19:49:56