Polars中计算百分比变化后的累积和异常问题求助
Polars计算pct_change后cum_sum出现inf的问题
我用Polars处理鸢尾花数据集时,想对sepal_width列先执行pct_change百分比变化计算,再计算其cum_sum累积和。但即使对无穷值和空值做了填充处理,结果仍出现inf无穷大,与预期不符。
我的代码
import polars as pl url = "https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv" df = pl.read_csv(url) change = ( df.with_columns(pl.col("sepal_width").shift(1, fill_value=0).pct_change(1).alias("pct_1")) .with_columns( pl.when(pl.col("pct_1").is_infinite()) .then(float(0)) .otherwise(pl.col("pct_1")) .fill_null(float(0)) .name.keep(), pl.col("pct_1").cum_sum().alias("cumsum_pct_1") ) )
当前输出
┌──────────────┬─────────────┬──────────────┬─────────────┬───────────┬───────────┬──────────────┐ │ sepal_length ┆ sepal_width ┆ petal_length ┆ petal_width ┆ species ┆ pct_1 ┆ cumsum_pct_1 │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ f64 ┆ f64 ┆ f64 ┆ f64 ┆ str ┆ f64 ┆ f64 │ ╞══════════════╪═════════════╪══════════════╪═════════════╪═══════════╪═══════════╪══════════════╡ │ 5.1 ┆ 3.5 ┆ 1.4 ┆ 0.2 ┆ setosa ┆ 0.0 ┆ null │ │ 4.9 ┆ 3.0 ┆ 1.4 ┆ 0.2 ┆ setosa ┆ 0.0 ┆ inf │ │ 4.7 ┆ 3.2 ┆ 1.3 ┆ 0.2 ┆ setosa ┆ -0.142857 ┆ inf │ │ 4.6 ┆ 3.1 ┆ 1.5 ┆ 0.2 ┆ setosa ┆ 0.066667 ┆ inf │
期望输出
┌──────────────┬─────────────┬──────────────┬─────────────┬───────────┬───────────┬──────────────┐ │ sepal_length ┆ sepal_width ┆ petal_length ┆ petal_width ┆ species ┆ pct_1 ┆ cumsum_pct_1 │ │ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- ┆ --- │ │ f64 ┆ f64 ┆ f64 ┆ f64 ┆ str ┆ f64 ┆ f64 │ ╞══════════════╪═════════════╪══════════════╪═════════════╪═══════════╪═══════════╪══════════════╡ │ 5.1 ┆ 3.5 ┆ 1.4 ┆ 0.2 ┆ setosa ┆ 0.0 ┆ 0.0 │ │ 4.9 ┆ 3.0 ┆ 1.4 ┆ 0.2 ┆ setosa ┆ 0.0 ┆ 0.0 │ │ 4.7 ┆ 3.2 ┆ 1.3 ┆ 0.2 ┆ setosa ┆ -0.142857 ┆ -0.142857 │ │ 4.6 ┆ 3.1 ┆ 1.5 ┆ 0.2 ┆ setosa ┆ 0.066667 ┆ -0.076190 │
问题原因及解决方法
问题根源
你当前代码的问题在于:同一with_columns方法中,所有表达式都是基于原始列计算的。也就是说,你在第二个with_columns里处理pct_1的逻辑,和计算cum_sum的逻辑是并行执行的,cum_sum引用的是还没被处理的原始pct_1列(该列包含inf值),所以累积和会变成inf。
修正方案
先处理好pct_1列,确保它没有inf和null,再基于处理后的列计算累积和。以下是两种简洁的修正代码:
方案一:分两步处理
import polars as pl url = "https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv" df = pl.read_csv(url) change = ( df.with_columns( # 先计算pct_change并直接清理inf和null,得到干净的pct_1 pl.col("sepal_width") .shift(1, fill_value=0) .pct_change(1) .fill_infinite(0.0) .fill_null(0.0) .alias("pct_1") ) .with_columns( # 基于处理后的pct_1计算累积和 pl.col("pct_1").cum_sum().alias("cumsum_pct_1") ) ) print(change.head())
方案二:同一表达式链完成计算
import polars as pl url = "https://raw.githubusercontent.com/uiuc-cse/data-fa14/gh-pages/data/iris.csv" df = pl.read_csv(url) change = df.with_columns( # 生成处理后的pct_1 pl.col("sepal_width") .shift(1, fill_value=0) .pct_change(1) .fill_infinite(0.0) .fill_null(0.0) .alias("pct_1"), # 直接在同一链中计算累积和 pl.col("sepal_width") .shift(1, fill_value=0) .pct_change(1) .fill_infinite(0.0) .fill_null(0.0) .cum_sum() .alias("cumsum_pct_1") ) print(change.head())
两种方案都能得到你期望的输出,核心是确保累积和计算的是已经清理过的pct_1列。
内容的提问来源于stack exchange,提问作者Hieu
相关产品推荐
相关产品推荐

