You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为60列数据集构建10×50的指定列相关矩阵?

生成指定列组间的10×50相关矩阵

下面提供Python和R两种常用数据分析环境的实现方法,直接针对目标列组计算相关系数,无需生成完整的60×60矩阵:

Python(Pandas)

假设你的数据集存储在Pandas DataFrame df 中,先明确两组目标列:

# 替换为你的10列列名列表
group1_cols = ['col_1', 'col_2', ..., 'col_10']
# 自动提取剩余50列
group2_cols = [col for col in df.columns if col not in group1_cols]

# 拆分数据集为两组
group1 = df[group1_cols]
group2 = df[group2_cols]

# 计算group1每列与group2所有列的皮尔逊相关系数,转置后得到10行50列矩阵
corr_matrix = group1.corrwith(group2).T

# 如需斯皮尔曼或肯德尔相关,添加method参数
# corr_matrix = group1.corrwith(group2, method='spearman').T

R语言

同样先定义两组列,用基础R即可快速实现:

# 替换为你的10列列名向量
group1_cols <- c("col_1", "col_2", ..., "col_10")
# 自动提取剩余50列
group2_cols <- setdiff(colnames(df), group1_cols)

# 拆分数据集
group1 <- df[, group1_cols]
group2 <- df[, group2_cols]

# 直接计算两组列的相关矩阵,结果为10×50维度
corr_matrix <- cor(group1, group2)

# 如需指定相关方法,添加method参数
# corr_matrix <- cor(group1, group2, method = "spearman")

这种方式仅计算目标列组间的相关系数,避免了冗余计算,在数据集较大时能显著节省内存和计算时间。

内容的提问来源于stack exchange,提问作者prof31

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 20:27:32