如何在DataFrame列的corr方法中指定Pearson相关系数计算方法?
解决办法及替代方案
直接指定method参数(最简便)
其实Series.corr()方法本身就支持传入method参数,你可以直接在调用时指定method='pearson',写法如下:
df_dataframe1['Food'].corr(df_dataframe1['Ingredient'], method='pearson')
注:pearson是默认计算方法,若只是明确指定默认值可省略该参数,但如果需要切换为spearman、kendall等其他方法,这个参数就会起到作用。
替代方案1:通过DataFrame生成相关矩阵提取结果
先选中目标列,调用DataFrame的corr()方法生成相关系数矩阵,再提取对应位置的数值:
# 生成两列的相关系数矩阵 corr_matrix = df_dataframe1[['Food', 'Ingredient']].corr(method='pearson') # 提取Food与Ingredient的相关系数 corr_value = corr_matrix.loc['Food', 'Ingredient']
替代方案2:使用scipy的pearsonr函数
如果需要同时获取相关系数和对应的显著性检验p值,可以用scipy.stats.pearsonr:
from scipy.stats import pearsonr corr_value, p_value = pearsonr(df_dataframe1['Food'], df_dataframe1['Ingredient'])
该方法会返回两个结果:第一个是皮尔逊相关系数,第二个是显著性检验的p值。
内容的提问来源于stack exchange,提问作者Alex Woolfe
相关产品推荐
相关产品推荐

