如何在Polars的group_by中基于另一列值聚合指定列求和
在Polars中按Weather分组后,基于Windy值聚合Price列的求和操作
你需要在按Weather分组的前提下,针对Windy的不同取值对Price列执行求和聚合,同时保留添加其他仅基于Weather分组的聚合操作的能力,不能直接用group_by(["Weather", "Windy"])。可以通过嵌套分组聚合实现需求:
实现代码
import polars as pl df = pl.DataFrame( data={ "Weather": ["Rain", "Sun", "Rain", "Sun", "Rain", "Sun", "Rain", "Sun"], "Price": [1, 2, 3, 4, 5, 6, 7, 8], "Windy": ["Y", "Y", "Y", "Y", "N", "N", "N", "N"] } ) df_agg = df.group_by("Weather").agg( # 将Windy和Price打包为struct,再嵌套分组求和 pl.struct(["Windy", "Price"]) .group_by("Windy") .agg(pl.col("Price").sum()) .alias("Price") ) print(df_agg)
运行结果
shape: (2, 2) ┌─────────┬────────────────────┐ │ Weather ┆ Price │ │ --- ┆ --- │ │ str ┆ list[struct[2]] │ ╞═════════╪════════════════════╡ │ Sun ┆ [{"Y",6},{"N",14}] │ │ Rain ┆ [{"Y",4},{"N",12}] │ └─────────┴────────────────────┘
原理说明
- 外层先按
Weather分组,确保可以同时添加其他仅基于Weather的聚合操作(比如直接追加pl.col("Price").sum().alias("Total_Price")) - 内层针对每个
Weather分组内的数据,先将Windy和Price打包为结构体,再按Windy分组对Price求和,最终得到每个Weather下不同Windy值对应的Price总和列表
内容的提问来源于stack exchange,提问作者Ynax
相关产品推荐
相关产品推荐

