Polars中GroupBy自定义众数函数及空值兼容问题求助
解决Polars分组求众数的窗口长度不匹配问题
当你尝试用Polars窗口函数计算分组众数时:
df.with_columns(pl.col("X").mode().over(['Y', 'Z']).name.prefix("mode_"))
因列X包含空值出现报错:
ComputeError: the length of the window expression did not match that of the group
本质原因是:若分组内X的所有值都是空值,mode()会返回空Series,窗口函数要求每个分组返回单个值,空Series无法匹配分组行数,导致报错。以下是两种可行解决方案:
方案一:自定义函数+分组聚合关联
适合需要复杂逻辑的场景,用Python自定义函数处理众数逻辑:
import polars as pl def custom_mode(x): filtered = x.drop_nulls() if not filtered: return None mode_result = filtered.mode() return mode_result[0] if mode_result else None # 先分组计算众数 mode_agg = df.group_by(['Y', 'Z']).agg( pl.col("X").map_elements(custom_mode).alias("mode_X") ) # 关联回原表保留所有行 result_df = df.join(mode_agg, on=['Y', 'Z'], how="left")
方案二:Polars内置函数优化(推荐)
用Polars原生方法替代Python UDF,性能更优,且无需手动处理空值逻辑:
result_df = df.with_columns( pl.col("X") .drop_nulls() # 过滤分组内的空值 .mode() # 计算众数 .first() # 取第一个众数,无众数时返回null .over(['Y', 'Z']) .alias("mode_X") )
这个方案中,drop_nulls()先移除分组内的空值,若分组内全为空则mode()返回空Series,first()会自动将空Series转为null,完美匹配窗口函数的长度要求,同时实现了"有众数返回众数,无则返回None"的需求。
内容的提问来源于stack exchange,提问作者Alessandro Togni
相关产品推荐
相关产品推荐

