You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars中对单列批量应用聚合操作(避免重复代码)

在Polars中对单列应用多聚合操作的简洁写法

当然有简洁写法,不用重复编写pl.col("列名")多次,以下是几种实用方案:

方案1:直接对列使用agg方法传入操作列表

Polars支持对列表达式直接调用agg方法,传入聚合操作的字符串列表,返回的结构列可以用unnest()展开为单独的列:

import polars as pl

df = pl.DataFrame({
    "a": [1, 3, 5, 7],
    "b": [2, 4, 6, 8]
})

# 对b列批量应用max、min、mean聚合
result = df.select(pl.col("b").agg(["max", "min", "mean"])).unnest("b")
print(result)

输出会自动生成b_max、b_min、b_mean三个列,对应各自的聚合结果。

方案2:列表推导式生成聚合表达式(灵活可控)

如果需要自定义列名,或者要对多列批量处理,用列表推导式更灵活:

# 定义需要的聚合操作
agg_ops = ["max", "min", "mean", "sum", "median"]
# 批量生成表达式,自定义别名
exprs = [pl.col("b").agg(op).alias(f"b_{op}") for op in agg_ops]
# 执行聚合
result = df.select(exprs)

这种方式可以轻松扩展到多列,比如同时处理a和b列:

cols_to_agg = ["a", "b"]
exprs = []
for col in cols_to_agg:
    exprs.extend([pl.col(col).agg(op).alias(f"{col}_{op}") for op in agg_ops])
result = df.select(exprs)

方案3:模拟Pandas的字典式聚合

如果需要对不同列指定不同的聚合操作,可以用字典来定义规则,再批量生成表达式:

# 定义各列对应的聚合操作
agg_spec = {
    "a": ["max", "median"],
    "b": ["mean", "min", "sum"]
}

exprs = []
for col, ops in agg_spec.items():
    exprs.extend([pl.col(col).agg(op).alias(f"{col}_{op}") for op in ops])

result = df.select(exprs)

这和Pandas中df.agg({"a": ["max"], "b": ["mean"]})的用法逻辑一致,完美替代重复编写列表达式的繁琐。

内容的提问来源于stack exchange,提问作者Mark Wang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 13:45:18