如何用Polars Rust按行计算指定前缀列的均值与标准差
解决方案
问题根源
你之前的代码错误在于对求和后的单值列调用std(1),这实际上是在计算整个列的标准差(按列聚合),而非每行10个EBIT值的行内标准差。
修正后的代码
利用Polars的list类型操作实现行内统计计算,同时通过通配符筛选目标列简化代码:
pub fn industry_beta_f(raw_data: DataFrame, marginal_tax_rate: Expr) -> DataFrame { let df = raw_data .lazy() // 将所有以"Sum of EBIT [CY"开头的列打包成每行的列表 .with_columns([cols("Sum of EBIT [CY*").list().alias("ebit_list")]) // 计算行内均值和标准差 .with_columns([ col("ebit_list").list.mean().alias("moments_mean"), // std的参数为ddof:1表示样本标准差,0表示总体标准差,按需调整 col("ebit_list").list.std(1).alias("moments_std"), ]) .with_columns([ when(col("moments_mean").gt(lit(0.0))) .then(col("moments_std") / col("moments_mean")) .otherwise(f64::NAN) .alias("Standard deviation in operating income (last 10 years)"), ]) .select([ col("Industry Name"), col("Number of firms"), col("Standard deviation in operating income (last 10 years)"), ]) .collect() .unwrap(); df }
关键改进点
- 通配符筛选列:用
cols("Sum of EBIT [CY*")自动匹配所有目标年份的EBIT列,避免手动重复编写10个列名。 - 行内统计计算:先将每行的EBIT值打包成列表列,再通过
list.mean()和list.std()实现行内的均值与标准差计算,完全符合需求。 - 避免冗余计算:无需手动求和再除以10,直接用
list.mean()更高效且易维护。
内容的提问来源于stack exchange,提问作者Carlos Arias
相关产品推荐
相关产品推荐

