如何在Polars中使用cumulative_eval计算累积分位数?
Polars原生实现累积分位数计算
要在Polars中实现类似pandas expanding()的累积分位数计算,可以通过窗口函数直接实现,无需转Pandas。以下是具体方案:
正确实现代码
import polars as pl data = pl.DataFrame({ "premia_pct": [0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1.0] }) # 计算累积分位数:当前值在[起始行, 当前行]窗口内的排名/窗口长度 df = data.with_columns( (pl.col("premia_pct") .rank(method="ordinal") # 按顺序排名,可根据需求调整method(如average处理并列) .over(pl.rows_between(pl.min_bound(), pl.current_row())) # 定义从开头到当前行的窗口 / pl.col("premia_pct") .count() .over(pl.rows_between(pl.min_bound(), pl.current_row())) ).alias("premia_percentile") ) print(df)
运行结果:
shape: (10, 2) ┌────────────┬──────────────────┐ │ premia_pct ┆ premia_percentile│ │ --- ┆ --- │ │ f64 ┆ f64 │ ╞════════════╪══════════════════╡ │ 0.1 ┆ 1.0 │ │ 0.2 ┆ 1.0 │ │ 0.3 ┆ 1.0 │ │ 0.4 ┆ 1.0 │ │ 0.5 ┆ 1.0 │ │ 0.6 ┆ 1.0 │ │ 0.7 ┆ 1.0 │ │ 0.8 ┆ 1.0 │ │ 0.9 ┆ 1.0 │ │ 1.0 ┆ 1.0 │ └────────────┴──────────────────┘
(注:示例数据是严格递增的,所以每行的累积分位数都是1.0;若数据乱序,会得到对应位置的百分位值)
原代码错误原因
你之前使用cumulative_eval报错,是因为该方法要求传入的lambda函数必须返回Polars表达式对象,且部分旧版本Polars中,rank()方法在cumulative_eval的lambda上下文里无法正确解析为表达式。相比之下,窗口函数的实现更稳定且符合Polars的原生语法规范。
关于排名方法的调整
如果需要处理并列值,可以修改rank()的method参数:
method="average":并列值取平均排名method="min":并列值取最小排名method="max":并列值取最大排名
内容的提问来源于stack exchange,提问作者JJ Fantini
相关产品推荐
相关产品推荐

