在Python Polars中实现按era分组计算特征与目标变量的相关性
在Polars中按分组计算特征与目标变量的相关性(对应Pandas的groupby+corrwith功能)
场景说明
你有如下Pandas DataFrame:
import pandas as pd d = {'era': ["a", "a", "b","b","c", "c"], 'feature1': [3, 4, 5, 6, 7, 8], 'feature2': [7, 8, 9, 10, 11, 12], 'target': [1, 2, 3, 4, 5 ,6]} df = pd.DataFrame(data=d)
在Pandas中,你通过以下代码按era分组,计算feature_cols = ['feature1', 'feature2']与TARGET_COL = 'target'的相关性:
feature_cols = ['feature1', 'feature2'] TARGET_COL = 'target' corrs_split = ( df .groupby("era") .apply(lambda d: d[feature_cols].corrwith(d[TARGET_COL])) )
该代码会生成以era为索引、特征名为列的结构化结果:
feature1 feature2 era a 1.0 1.0 b 1.0 1.0 c 1.0 1.0
Polars实现方案
以下两种方式可以在Polars中得到完全一致的结构化结果:
方式一:使用map_groups
通过分组映射,在每个组内单独计算特征与目标的相关性,再拼接分组标识:
import polars as pl # 构建Polars DataFrame df_pl = pl.DataFrame(d) feature_cols = ['feature1', 'feature2'] TARGET_COL = 'target' corrs_split_pl = ( df_pl .group_by("era") .map_groups( lambda group: pl.DataFrame( {col: [group[col].corr(group[TARGET_COL])] for col in feature_cols} ).with_columns(era=pl.lit(group["era"].unique().item())) ) .select("era", *feature_cols) ) print(corrs_split_pl)
输出结果:
shape: (3, 3) ┌─────┬──────────┬──────────┐ │ era ┆ feature1 ┆ feature2 │ │ --- ┆ --- ┆ --- │ │ str ┆ f64 ┆ f64 │ ╞═════╪══════════╪══════════╡ │ a ┆ 1.0 ┆ 1.0 │ │ b ┆ 1.0 ┆ 1.0 │ │ c ┆ 1.0 ┆ 1.0 │ └─────┴──────────┴──────────┘
方式二:使用melt+pivot
先将特征列转成长格式,分组计算相关性后再转回宽格式,逻辑更简洁:
corrs_split_pl = ( df_pl .melt(id_vars=["era", TARGET_COL], value_vars=feature_cols, variable_name="feature") .group_by(["era", "feature"]) .agg(correlation=pl.corr("value", TARGET_COL)) .pivot(index="era", columns="feature", values="correlation") ) print(corrs_split_pl)
该代码会输出和方式一完全相同的结构化结果。
内容的提问来源于stack exchange,提问作者jbssm
相关产品推荐
相关产品推荐

