You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中按百分比差值对数据项进行分组?

按百分比阈值分组并标记组首元素

要实现将数值按「元素与当前组首元素的差值不超过指定百分比」的规则分组,同时生成组编号和组首元素列,可以用Polars的fold函数完成带状态的累积计算,以下是具体实现:

代码实现

import polars as pl

# 原始数据
df = pl.DataFrame({'a': [100, 103, 105, 106, 105, 104, 103, 106, 100, 102]})

# 设置百分比阈值(此处为5%)
threshold = 0.05

# 计算分组和组首值
result_df = df.with_columns(
    # 使用fold进行累积状态计算,跟踪当前组号和组首值
    pl.fold(
        # 初始状态:组号0,组首值为第一行的a值
        acc=pl.struct(group=pl.lit(0), first_val=pl.col('a').first()),
        # 每一行的判断逻辑
        function=lambda acc, x: pl.when(
            x > acc['first_val'] * (1 + threshold)
        ).then(
            # 超过阈值,新建组:组号+1,组首值更新为当前x
            pl.struct(group=acc['group'] + 1, first_val=x)
        ).otherwise(
            # 未超过阈值,保持当前组状态
            pl.struct(group=acc['group'], first_val=acc['first_val'])
        ),
        exprs=pl.col('a')
    ).alias('group_state')
).with_columns(
    # 从结构体中提取组号和组首值列
    pl.col('group_state').struct.field('group').alias('group'),
    pl.col('group_state').struct.field('first_val').alias('groupFirst')
).drop('group_state')  # 删除临时状态列

print(result_df)

输出结果

shape: (10, 3)
┌─────┬───────┬────────────┐
│ a   ┆ group ┆ groupFirst │
│ --- ┆ ---   ┆ ---        │
│ i64 ┆ i64   ┆ i64        │
╞═════╪═══════╪════════════╡
│ 100 ┆ 0     ┆ 100        │
│ 103 ┆ 0     ┆ 100        │
│ 105 ┆ 1     ┆ 105        │
│ 106 ┆ 1     ┆ 105        │
│ 105 ┆ 1     ┆ 105        │
│ 104 ┆ 1     ┆ 105        │
│ 103 ┆ 1     ┆ 105        │
│ 106 ┆ 1     ┆ 105        │
│ 100 ┆ 2     ┆ 100        │
│ 102 ┆ 2     ┆ 100        │
└─────┴───────┴────────────┘

逻辑说明

  • pl.fold是Polars中用于带状态累积计算的函数,这里用它跟踪两个状态:当前组的编号group和组首元素first_val
  • 初始状态取第一行的a值作为第一个组的组首,组号从0开始
  • 对每一行的a值进行判断:如果当前值超过组首值的(1+阈值)倍(即高出5%),则触发新组创建,更新组号和组首值;否则保持当前组状态
  • 最后将结构体中的状态字段拆分,得到目标列

内容的提问来源于stack exchange,提问作者aiv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.01 06:25:33