You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Polars分组聚合将分类值计数转独立列并计算总计的方法

Polars 分组后将分类值拆为独立计数列实现方案

你需要的是分组后将分类列的取值转为独立列统计频次,再计算总和,有两种简洁实现方式:


方案1:使用pivot透视函数(最适配场景)

pivot是专门处理行转列统计的函数,不需要额外做长表转宽表的二次处理,直接一步完成分组计数:

import polars as pl

df = pl.DataFrame({
    'ID': [0, 0,1, 1, 1, 0], 
    'Type': ['Fire', 'Fire', 'Fire', 'Water', 'Water', 'Water'], 
})

result = df.pivot(
    values="Type",
    index="ID",
    columns="Type",
    aggregate_function="count"
).with_columns(
    total = pl.col("Fire") + pl.col("Water")
# 如果需要列名和示例一致为全小写,加下面这行
# ).rename({"Fire":"fire", "Water":"water"})
)

print(result)

参数说明:

  • index="ID":指定ID作为分组维度,作为结果的行标识
  • columns="Type":将Type列的唯一取值(Fire、Water)拆分为独立列
  • aggregate_function="count":对每个分组下的分类值做计数统计
  • 后续通过with_columns新增total列,直接求和两个分类的计数即可得到总条数

方案2:group_by聚合时直接条件计数

如果你更习惯group_by的写法,可以在聚合阶段直接按条件过滤计数,不需要生成中间长表结果:

result = df.group_by("ID").agg(
    fire = pl.col("Type").filter(pl.col("Type") == "Fire").count(),
    water = pl.col("Type").filter(pl.col("Type") == "Water").count()
).with_columns(
    total = pl.col("fire") + pl.col("water")
)

两种方法最终输出的结果和你给出的预期结构完全一致:

shape: (2, 4)
┌─────┬──────┬───────┬───────┐
│ ID  ┆ fire ┆ water ┆ total │
│ --- ┆ ---  ┆ ---   ┆ ---   │
│ i64 ┆ u32  ┆ u32   ┆ u32   │
╞═════╪══════╪═══════╪═══════╡
│ 0   ┆ 2    ┆ 1     ┆ 3     │
│ 1   ┆ 1    ┆ 2     ┆ 3     │
└─────┴──────┴───────┴───────┘

内容的提问来源于stack exchange,提问作者supersick

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 11:36:21