You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中不折叠DataFrame完成分组级计算与取值?

问题描述

我希望在不创建第二个DataFrame的前提下,完成分组级别的统计计算。目前我的做法是先生成包含所需聚合结果的第二个DataFrame,再将其合并回原DataFrame。

举个简单示例:

import polars as pl

df = pl.DataFrame( {'name' : ['Steve', 'Larry', 'Tom', 'Steve', 'Tom', 'Steve'],
                     'points': range(6)})
print(df)

输出:

shape: (6, 2)
┌───────┬────────┐
│ name  ┆ points │
│ ---   ┆ ---    │
│ str   ┆ i64    │
╞═══════╪════════╡
│ Steve ┆ 0      │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┤
│ Larry ┆ 1      │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┤
│ Tom   ┆ 2      │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┤
│ Steve ┆ 3      │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┤
│ Tom   ┆ 4      │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┤
│ Steve ┆ 5      │
└───────┴────────┘

第二步,生成记录每个分组大小的额外DataFrame:

entries= df.groupby('name').agg(pl.count().alias('entries'))
print(entries)

输出:

shape: (3, 2)
┌───────┬─────────┐
│ name  ┆ entries │
│ ---   ┆ ---     │
│ str   ┆ u32     │
╞═══════╪═════════╡
│ Steve ┆ 3       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Tom   ┆ 2       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Larry ┆ 1       │
└───────┴─────────┘

第三步,合并回原DataFrame:

print(df.join(entries, left_on='name', right_on='name', how='left'))

输出:

shape: (6, 3)
┌───────┬────────┬─────────┐
│ name  ┆ points ┆ entries │
│ ---   ┆ ---    ┆ ---     │
│ str   ┆ i64    ┆ u32     │
╞═══════╪════════╪═════════╡
│ Steve ┆ 0      ┆ 3       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Larry ┆ 1      ┆ 1       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Tom   ┆ 2      ┆ 2       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Steve ┆ 3      ┆ 3       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Tom   ┆ 4      ┆ 2       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Steve ┆ 5      ┆ 3       │
└───────┴────────┴─────────┘

是否有办法避免这种分步操作?我感觉可以使用over窗口函数,但尚未找到具体实现方法。


解决方案

完全可以用Polars的over窗口函数一步完成,不需要额外创建DataFrame再合并。直接在with_columns中使用聚合函数搭配over('分组列')即可,这样会自动将分组统计结果广播到原DataFrame的每一行。

针对你的示例,代码如下:

import polars as pl

df = pl.DataFrame( {'name' : ['Steve', 'Larry', 'Tom', 'Steve', 'Tom', 'Steve'],
                     'points': range(6)})

# 一步添加分组条目数字段
result = df.with_columns(
    pl.count().over('name').alias('entries')
)

print(result)

输出结果和你之前分步操作的完全一致:

shape: (6, 3)
┌───────┬────────┬─────────┐
│ name  ┆ points ┆ entries │
│ ---   ┆ ---    ┆ ---     │
│ str   ┆ i64    ┆ u32     │
╞═══════╪════════╪═════════╡
│ Steve ┆ 0      ┆ 3       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Larry ┆ 1      ┆ 1       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Tom   ┆ 2      ┆ 2       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Steve ┆ 3      ┆ 3       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Tom   ┆ 4      ┆ 2       │
├╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌┼╌╌╌╌╌╌╌╌╌┤
│ Steve ┆ 5      ┆ 3       │
└───────┴────────┴─────────┘

扩展说明

除了pl.count(),其他聚合函数比如pl.sum()、pl.mean()、pl.max()等都可以搭配over使用,实现不同的分组统计需求。比如要添加每个分组的points总和:

result = df.with_columns(
    pl.count().over('name').alias('entries'),
    pl.sum('points').over('name').alias('total_points')
)

这样就能一次性添加多个分组统计字段,全程不需要额外创建中间DataFrame。

内容的提问来源于stack exchange,提问作者pedrosaurio

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.20 11:27:25