You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何按分组保留Polars DataFrame中的奇数位置行?

按分组保留Polars DataFrame的奇数位置行

问题描述

现有一个已排序的Polars DataFrame(版本0.19),包含分组列group_col,需要按分组筛选,保留每组的奇数位置行(即删除每组的第2、4、...行,保留第1、3、...行)。已知DataFrame总行数为偶数。

示例输入:

┌───────────┬───────────┐
│ group_col ┆ value_col │
│ ---       ┆ ---       │
│ i64       ┆ i64       │
╞═══════════╪═══════════╡
│ 1         ┆ 10        │
│ 1         ┆ 20        │
│ 1         ┆ 30        │
│ 1         ┆ 40        │
│ 2         ┆ 50        │
│ 2         ┆ 60        │
│ 3         ┆ 70        │
│ 3         ┆ 80        │
└───────────┴───────────┘

示例输出:

┌───────────┬───────────┐
│ group_col ┆ value_col │
│ ---       ┆ ---       │
│ i64       ┆ i64       │
╞═══════════╪═══════════╡
│ 1         ┆ 10        │
│ 1         ┆ 30        │
│ 2         ┆ 50        │
│ 3         ┆ 70        │
└───────────┴───────────┘

原DataFrame定义代码:

import polars as pl

df = pl.DataFrame({
    'group_col': [1, 1, 1, 1, 2, 2, 3, 3],
    'value_col': [10, 20, 30, 40, 50, 60, 70, 80]
})

解决方案

利用Polars的窗口函数over()生成分组内的行计数,再筛选符合条件的行:

# 生成分组内行计数,筛选奇数位置行(从1开始数),移除辅助列
result = df.with_columns(
    row_num=pl.cum_count().over('group_col')
).filter(pl.col('row_num') % 2 == 0).drop('row_num')

# 打印结果
print(result)

代码解释

  1. pl.cum_count().over('group_col'):为每个分组内的行生成从0开始的连续计数(例如group1的行计数为[0,1,2,3])。
  2. filter(pl.col('row_num') % 2 == 0):保留计数为偶数的行,对应分组内第1、3、5...行(从1开始计数的奇数位置)。
  3. drop('row_num'):移除用于筛选的辅助列,还原原DataFrame的列结构。

替代实现(分组聚合+展开)

也可以通过分组聚合后展开的方式实现,效果一致:

result = df.group_by('group_col', maintain_order=True).agg(
    pl.all().filter(pl.cum_count() % 2 == 0)
).explode(pl.all())

这里maintain_order=True确保分组顺序与原DataFrame一致,agg()内对每列筛选符合条件的行,最后用explode()将聚合后的列表展开为行。

内容的提问来源于stack exchange,提问作者MantleMan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 09:00:22