如何在Polars分组计算均值时忽略0.0值?
Polars分组求均值时忽略0.0的解决办法
方案一:把0.0替换成NaN再计算均值
Polars的mean()函数会自动忽略NaN值,所以先把所有数值列里的0.0换成None(Polars会自动转成对应类型的NaN)就行。
批量替换数值列的0.0
先筛选出需要处理的列(排除ts和month这类非数值列),然后用两种方式实现替换:
import polars as pl # 筛选出所有要处理的数值列 numeric_cols = [col for col in df.columns if col not in ["ts", "month"]] # 方法A:用replace直接替换0.0为None df_processed = df.with_columns( pl.col(numeric_cols).replace(0.0, None) ) # 方法B:用条件表达式(适合需要判断精度的场景) df_processed = df.with_columns( pl.col(numeric_cols).when(pl.col() != 0.0).otherwise(None) ) # 分组求均值 result = df_processed.group_by("month", maintain_order=True).mean()
方案二:聚合时直接过滤0.0(不用提前改数据)
更高效的方式是直接在聚合阶段排除0.0,不用修改原DataFrame:
import polars as pl numeric_cols = [col for col in df.columns if col not in ["ts", "month"]] result = df.group_by("month", maintain_order=True).agg( # 对每个数值列,先过滤掉0.0再求均值 pl.col(col).filter(pl.col() != 0.0).mean().alias(col) for col in numeric_cols )
这种方式不会改动原始数据,逻辑也更直接,推荐优先使用。
额外提示
如果你的数据有浮点精度问题(比如存在类似1e-17这种极接近0的数),可以用绝对值判断来避免误判:
# 替换时的精度判断 pl.col(numeric_cols).when(pl.col().abs() > 1e-9).otherwise(None) # 聚合时的精度判断 pl.col(col).filter(pl.col().abs() > 1e-9).mean().alias(col)
内容的提问来源于stack exchange,提问作者Glenn Pierce
相关产品推荐
相关产品推荐

