Polars中rolling_map自定义函数TypeError问题(当前版本已修复)
问题解决:Polars rolling_map 自定义函数报错处理
更新说明
该问题目前已在Polars中解决,原函数可正常运行无报错。
原问题描述
希望在Polars的rolling_map中使用自定义函数,但执行代码时遇到TypeError错误。
原代码
def ts_rank(expr: pl.Expr, window: int) -> pl.Expr: res = expr.cast(pl.Float64).rolling_map( lambda s: s.rank(method='average', descending=False)[-1]/s.is_not_null().sum(), window_size = window, min_periods = window//2).over('a') return res df = pl.DataFrame({"a": [1, 1, 1, 1, 2, 2, 2, 2], "b": [None, None, None, 1, 4, 2, 3, 8]}) df.with_columns(ts_rank(pl.col('b'),4).alias('rank'))
报错信息
PanicException: python function failed: PyErr { type: , value: TypeError("unsupported operand type(s) for /: 'NoneType' and 'int'"), traceback: Some() }
问题:这是否是实现rolling_rank的正确Polars方式?(出于需求,必须以Expr形式实现,不能使用DataFrame.rolling)
问题分析与解决
报错原因
当滑动窗口内所有值都是None时,s.rank(...)返回的序列最后一个元素是None,此时执行除法运算就会触发类型错误。比如分组a=1的前3行,窗口内全为None,s.rank()结果全是None,取[-1]得到None,进而导致None / 整数的非法运算。
兼容旧版本的修复方案
在自定义lambda中加入空值判断,当窗口内无有效数据时直接返回None,避免运算错误:
def ts_rank(expr: pl.Expr, window: int) -> pl.Expr: res = expr.cast(pl.Float64).rolling_map( lambda s: s.rank(method='average', descending=False)[-1]/s.is_not_null().sum() if s.is_not_null().sum() > 0 else None, window_size=window, min_periods=window//2 ).over('a') return res
实现方式合理性
用rolling_map结合.over()分组的方式,确实是符合需求的Expr形式实现rolling_rank的正确思路。Polars后续版本优化了空值处理逻辑,所以你的原函数现在能正常运行,但在旧版本中需要加入上述空值判断来规避报错。
内容的提问来源于stack exchange,提问作者Jun
相关产品推荐
相关产品推荐

