如何将自定义非中心化Pearson相关函数应用于DataFrame列对生成相关矩阵
生成非中心化Pearson相关矩阵的实现方法
数据背景
现有包含日期(2022年9月30日至2022年11月30日,不含周末)及15只股票价格的DataFrame,示例数据如下:
| DATES | A | B | C | D | E |
|---|---|---|---|---|---|
| 30/09/22 | 100.5 | 151.3 | 233.4 | 237.2 | 38.42 |
| 01/10/22 | 101.5 | 148.0 | 237.6 | 232.2 | 38.54 |
| 02/10/22 | 102.2 | 147.6 | 238.3 | 231.4 | 39.32 |
| 03/10/22 | 103.4 | 145.7 | 239.2 | 232.2 | 39.54 |
已实现的中心化相关矩阵代码
已通过以下代码生成中心化Pearson相关矩阵:
import pandas as pd import numpy as np df = pd.read_excel(file_path, sheet_name) df = df.dropna() # 移除存在缺失股价的日期 log_df = df.set_index("DATES").pipe(lambda d: np.log(d.div(d.shift()))).reset_index() corrM = log_df.corr()
自定义非中心化相关系数函数
已定义非中心化Pearson相关系数计算函数:
def uncentered_correlation(x, y): x_dim = len(x) y_dim = len(y) xy = 0 xx = 0 yy = 0 for i in range(x_dim): xy += x[i] * y[i] xx += x[i] ** 2.0 yy += y[i] ** 2.0 corr = xy / np.sqrt(xx * yy) return corr
应用函数生成非中心化相关矩阵
方法一:直接调用pandas的corr方法
pandas的DataFrame.corr()支持传入自定义相关系数函数,一步生成对称矩阵:
# 提取对数收益率列(排除日期列) return_cols = log_df.columns.drop("DATES") # 生成非中心化相关矩阵 uncentered_corrM = log_df[return_cols].corr(method=uncentered_correlation)
方法二:手动遍历列对构建矩阵
若需更灵活控制,可手动遍历所有列对填充矩阵:
import itertools cols = return_cols # 初始化空矩阵 uncentered_corrM = pd.DataFrame(index=cols, columns=cols) # 遍历所有列组合(含自身) for col1, col2 in itertools.product(cols, cols): uncentered_corrM.loc[col1, col2] = uncentered_correlation(log_df[col1], log_df[col2]) # 转换为浮点型(可选) uncentered_corrM = uncentered_corrM.astype(float)
函数优化建议
原循环实现效率较低,改用numpy向量运算可大幅提升速度:
def uncentered_correlation(x, y): xy = np.dot(x, y) xx = np.dot(x, x) yy = np.dot(y, y) return xy / np.sqrt(xx * yy)
内容的提问来源于stack exchange,提问作者Guillem
相关产品推荐
相关产品推荐

