You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars分组DataFrame中创建非聚合变量

Polars按组计算累积和的正确实现

初始数据集代码

import polars as pl

df_a = pl.DataFrame({'recipe': 'A','values': [1,2,3]})
df_b = pl.DataFrame({'recipe': 'B','values': [1,2,3]})

df = pl.concat([df_a, df_b], rechunk=True)

问题场景

想要按recipe分组计算values的累积和,但直接使用group_by后调用with_columns的写法无法运行:

df.group_by('recipe').with_columns(cum_val = np.cumsum(pl.col('values')))

对比R中tidyverse的实现方式:

library(tidyverse)

mtcars |> group_by(cyl) |> mutate(cum_wt=cumsum(wt))

正确实现方式

方法1:窗口函数(推荐,和R逻辑一致)

Polars中可以直接用窗口函数over()指定分组键,无需提前group_by,直接对每行计算分组内的累积和:

df.with_columns(
    cum_val = pl.col('values').cumsum().over('recipe')
)

执行后结果如下:

shape: (6, 3)
┌─────────┬────────┬────────┐
│ recipe  ┆ values ┆ cum_val│
│ ---     ┆ ---    ┆ ---    │
│ str     ┆ i64    ┆ i64    │
╞═════════╪════════╪════════╡
│ A       ┆ 1      ┆ 1      │
│ A       ┆ 2      ┆ 3      │
│ A       ┆ 3      ┆ 6      │
│ B       ┆ 1      ┆ 1      │
│ B       ┆ 2      ┆ 3      │
│ B       ┆ 3      ┆ 6      │
└─────────┴────────┴────────┘

方法2:分组聚合后展开(不推荐,效率较低)

如果一定要使用group_by,可以先分组生成每组的累积和列表,再通过explode展开后和原表关联:

# 分组聚合生成每组的累积和
grouped_cum = df.group_by('recipe').agg(
    cum_val = pl.col('values').cumsum()
)
# 展开列表并关联原表
result = df.join(grouped_cum.explode('cum_val'), on='recipe')

原代码报错原因

Polars中group_by后的with_columns是针对分组聚合结果添加列,而非对原数据的每行进行分组计算,因此无法直接实现类似R中group_by + mutate的逐行分组逻辑,必须使用窗口函数来指定分组范围。

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 12:22:40