You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Polars分组计算均值时忽略0.0值?

Polars分组求均值时忽略0.0的解决办法

方案一:把0.0替换成NaN再计算均值

Polars的mean()函数会自动忽略NaN值,所以先把所有数值列里的0.0换成None(Polars会自动转成对应类型的NaN)就行。

批量替换数值列的0.0

先筛选出需要处理的列(排除ts和month这类非数值列),然后用两种方式实现替换:

import polars as pl

# 筛选出所有要处理的数值列
numeric_cols = [col for col in df.columns if col not in ["ts", "month"]]

# 方法A:用replace直接替换0.0为None
df_processed = df.with_columns(
    pl.col(numeric_cols).replace(0.0, None)
)

# 方法B:用条件表达式(适合需要判断精度的场景)
df_processed = df.with_columns(
    pl.col(numeric_cols).when(pl.col() != 0.0).otherwise(None)
)

# 分组求均值
result = df_processed.group_by("month", maintain_order=True).mean()

方案二:聚合时直接过滤0.0(不用提前改数据)

更高效的方式是直接在聚合阶段排除0.0,不用修改原DataFrame:

import polars as pl

numeric_cols = [col for col in df.columns if col not in ["ts", "month"]]

result = df.group_by("month", maintain_order=True).agg(
    # 对每个数值列,先过滤掉0.0再求均值
    pl.col(col).filter(pl.col() != 0.0).mean().alias(col)
    for col in numeric_cols
)

这种方式不会改动原始数据,逻辑也更直接,推荐优先使用。

额外提示

如果你的数据有浮点精度问题(比如存在类似1e-17这种极接近0的数),可以用绝对值判断来避免误判:

# 替换时的精度判断
pl.col(numeric_cols).when(pl.col().abs() > 1e-9).otherwise(None)

# 聚合时的精度判断
pl.col(col).filter(pl.col().abs() > 1e-9).mean().alias(col)

内容的提问来源于stack exchange,提问作者Glenn Pierce

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 07:53:30