按分组计算变量Spearman相关性时结果重复的问题解决
问题描述
基于group、meeting、session_phase三个分组变量,计算var1至var5两两之间的Spearman相关性并生成新数据框,要求忽略NA值。但运行以下代码后,所有分组的相关系数结果完全相同,不符合每个分组对应独立结果的预期:
library(dplyr) calc_cor <- function(df) { list( Cor1x2 = cor(df$var1, df$var2, method = "spearman", use = "pairwise.complete.obs"), Cor1x3 = cor(df$var1, df$var3, method = "spearman", use = "pairwise.complete.obs"), Cor1x4 = cor(df$var1, df$var4, method = "spearman", use = "pairwise.complete.obs"), Cor1x5 = cor(df$var1, df$var5, method = "spearman", use = "pairwise.complete.obs"), Cor2x3 = cor(df$var2, df$var3, method = "spearman", use = "pairwise.complete.obs"), Cor2x4 = cor(df$var2, df$var4, method = "spearman", use = "pairwise.complete.obs"), Cor2x5 = cor(df$var2, df$var5, method = "spearman", use = "pairwise.complete.obs"), Cor3x4 = cor(df$var3, df$var4, method = "spearman", use = "pairwise.complete.obs"), Cor3x5 = cor(df$var3, df$var5, method = "spearman", use = "pairwise.complete.obs"), Cor4x5 = cor(df$var4, df$var5, method = "spearman", use = "pairwise.complete.obs") ) } result <- df %>% group_by(group, meeting, session_phase) %>% summarise(correlations = list(calc_cor(.))) %>% unnest_wider(correlations) %>% as.data.frame()
输入数据中var4和var5存在大量NA值,分组变量group取值1-55,meeting取值1-6,session_phase取值1-2。
原因分析
- 函数返回格式与
summarise适配问题:原代码中calc_cor返回列表,再用list()包裹后传入summarise,导致unnest_wider未正确解析分组后的独立结果,反而复用了全局数据的计算值。 - 分组观测数不足:若部分分组的有效观测数≤1,
cor函数会返回固定值(如1或NA),导致多个分组结果看似一致。
修正方案
方案1:调整函数返回格式与调用逻辑
将calc_cor改为返回tibble,直接在summarise中展开结果,无需额外嵌套与反嵌套操作:
library(dplyr) calc_cor <- function(df) { tibble( Cor1x2 = cor(df$var1, df$var2, method = "spearman", use = "pairwise.complete.obs"), Cor1x3 = cor(df$var1, df$var3, method = "spearman", use = "pairwise.complete.obs"), Cor1x4 = cor(df$var1, df$var4, method = "spearman", use = "pairwise.complete.obs"), Cor1x5 = cor(df$var1, df$var5, method = "spearman", use = "pairwise.complete.obs"), Cor2x3 = cor(df$var2, df$var3, method = "spearman", use = "pairwise.complete.obs"), Cor2x4 = cor(df$var2, df$var4, method = "spearman", use = "pairwise.complete.obs"), Cor2x5 = cor(df$var2, df$var5, method = "spearman", use = "pairwise.complete.obs"), Cor3x4 = cor(df$var3, df$var4, method = "spearman", use = "pairwise.complete.obs"), Cor3x5 = cor(df$var3, df$var5, method = "spearman", use = "pairwise.complete.obs"), Cor4x5 = cor(df$var4, df$var5, method = "spearman", use = "pairwise.complete.obs") ) } # 计算分组相关性,同时过滤观测数不足2的分组 result <- df %>% group_by(group, meeting, session_phase) %>% filter(n() >= 2) %>% # 至少2个观测才能计算有效相关性 summarise(calc_cor(.), .groups = "drop") %>% as.data.frame()
方案2:简化相关性计算逻辑(可选)
通过矩阵计算批量生成两两相关性,代码更简洁且不易出错:
library(dplyr) library(tidyr) calc_cor_matrix <- function(df) { # 提取var1-var5列,计算Spearman相关矩阵 cor_mat <- cor(df %>% select(starts_with("var")), method = "spearman", use = "pairwise.complete.obs") # 提取上三角的两两相关值,排除对角线 cor_pairs <- cor_mat[upper.tri(cor_mat)] # 生成对应的列名 pair_names <- combn(paste0("var", 1:5), 2, function(x) paste0("Cor", x[1], "x", x[2])) # 返回tibble格式结果 as_tibble(setNames(list(cor_pairs), pair_names)) } result <- df %>% group_by(group, meeting, session_phase) %>% filter(n() >= 2) %>% summarise(calc_cor_matrix(.), .groups = "drop") %>% as.data.frame()
验证步骤
- 先检查每个分组的观测数,确认是否存在大量样本量不足的情况:
df %>% group_by(group, meeting, session_phase) %>% tally(name = "obs_count") %>% arrange(obs_count)
- 运行修正后的代码,随机抽取几个分组的结果,验证相关性是否存在差异。
内容的提问来源于stack exchange,提问作者Jan
相关产品推荐
相关产品推荐

