能否用Polars单链式调用实现计数与百分比计算?
Polars单链式调用实现分组计数与百分比计算
问题
观察到Polars的其他示例中多数操作可通过单链式调用完成,请问能否用单链式调用实现下述示例的计数与百分比计算?是否存在简化方案?
原代码:
import polars as pl scores = pl.DataFrame({ 'zone': ['North', 'North', 'North', 'South', 'East', 'East', 'East', 'East'], 'score': [78, 39, 76, 56, 67, 89, 100, 55] }) cnt = scores.group_by("zone").len() cnt.with_columns( (100 * pl.col("len") / pl.col("len").sum()) .round(2) .cast(str) .str.replace(r"$", "%") .alias("perc") )
原输出:
shape: (3, 3) ┌───────┬─────┬───────┐ │ zone ┆ len ┆ perc │ │ --- ┆ --- ┆ --- │ │ str ┆ u32 ┆ str │ ╞═══════╪═════╪═══════╡ │ South ┆ 1 ┆ 12.5% │ │ East ┆ 4 ┆ 50.0% │ │ North ┆ 3 ┆ 37.5% │ └───────┴─────┴───────┘
解答
当然可以用单链式调用实现,而且有更简洁的简化方案:
1. 基础单链式调用
直接去掉中间变量cnt,把两步操作合并为一条链式调用:
import polars as pl scores = pl.DataFrame({ 'zone': ['North', 'North', 'North', 'South', 'East', 'East', 'East', 'East'], 'score': [78, 39, 76, 56, 67, 89, 100, 55] }) ( scores .group_by("zone") .len() .with_columns( (100 * pl.col("len") / pl.col("len").sum()) .round(2) .cast(str) .str.replace(r"$", "%") .alias("perc") ) )
2. 简化百分比格式化
原代码中round+cast+str.replace的组合可以用pl.format函数一步完成,代码更简洁易读:
( scores .group_by("zone") .len() .with_columns( pl.format("{:.2f}%", 100 * pl.col("len") / pl.col("len").sum()).alias("perc") ) )
pl.format会自动完成四舍五入、数值转字符串、拼接百分号的操作,效果和原代码完全一致。
3. 更紧凑的聚合写法
如果想把计数和百分比计算放在同一个聚合步骤里,可以直接用agg方法:
( scores .group_by("zone") .agg( len=pl.len(), perc=pl.format("{:.2f}%", 100 * pl.len() / pl.len().sum()) ) )
这种写法把所有计算逻辑都封装在分组聚合阶段,也是单链式调用的一种高效写法。
内容的提问来源于stack exchange,提问作者Vincent
相关产品推荐
相关产品推荐

