如何在Pandas中为多列获取p值与Pearson相关系数r
用scipy.stats生成带相关系数和p值的多层索引矩阵
需求说明
基于给定的DataFrame,生成一个多层索引矩阵:行维度为「变量名+统计量类型(r表示相关系数,p表示p值)」,列维度为变量名,矩阵内容对应两两变量的相关系数及对应的p值,同时适配新版scipy.stats的API(旧版直接返回(r,p)元组,新版返回结果对象)。
示例数据
import pandas as pd from scipy.stats import pearsonr # 构造示例DataFrame x = pd.DataFrame( list( zip( [1,2,3,4,5,6], [5, 7, 8, 4, 2, 8], [13, 16, 12, 11, 9, 10] ) ), columns=['a', 'b', 'c'] )
问题分析
之前的for循环仅处理相邻列,无法覆盖所有变量两两组合;且新版scipy.stats的统计函数(如pearsonr、spearmanr)不再直接返回(r,p)元组,而是返回包含statistic(相关系数)和pvalue(p值)属性的结果对象,需要调整取值方式。
解决方案
实现步骤
- 初始化空矩阵存储相关系数和p值
- 遍历所有变量两两组合,计算并填充值
- 将两个矩阵合并为多层索引结构
完整代码
# 获取所有列名 cols = x.columns # 初始化相关系数矩阵和p值矩阵 r_matrix = pd.DataFrame(index=cols, columns=cols) p_matrix = pd.DataFrame(index=cols, columns=cols) # 遍历所有变量对,填充矩阵 for row_var in cols: for col_var in cols: if row_var == col_var: # 变量与自身的相关系数为1,p值设为0 r_matrix.loc[row_var, col_var] = 1.0 p_matrix.loc[row_var, col_var] = 0.0 else: # 新版scipy调用方式:通过结果对象的statistic和pvalue属性取值 result = pearsonr(x[row_var], x[col_var]) # 保留两位小数,可根据需求调整 r_matrix.loc[row_var, col_var] = round(result.statistic, 2) p_matrix.loc[row_var, col_var] = round(result.pvalue, 2) # 为矩阵添加统计量类型的层级索引 r_matrix = r_matrix.assign(stat_type='r').set_index('stat_type', append=True) p_matrix = p_matrix.assign(stat_type='p').set_index('stat_type', append=True) # 合并两个矩阵并按变量名排序 final_result = pd.concat([r_matrix, p_matrix]).sort_index(level=0) # 打印结果 print(final_result)
输出结果
| a | b | c | ||
|---|---|---|---|---|
| a | r | 1.0 | -0.09 | -0.80 |
| p | 0.00 | 0.87 | 0.06 | |
| b | r | -0.09 | 1.0 | 0.42 |
| p | 0.87 | 0.00 | 0.41 | |
| c | r | -0.80 | 0.42 | 1.0 |
| p | 0.06 | 0.41 | 0.00 |
适配spearmanr的说明
如果需要使用斯皮尔曼相关系数,只需将代码中的pearsonr替换为spearmanr即可,新版API的调用方式一致,均通过statistic和pvalue属性取值。
内容的提问来源于stack exchange,提问作者KevOMalley743
相关产品推荐
相关产品推荐

