You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按分组计算变量Spearman相关性时结果重复的问题解决

问题描述

基于group、meeting、session_phase三个分组变量,计算var1至var5两两之间的Spearman相关性并生成新数据框,要求忽略NA值。但运行以下代码后,所有分组的相关系数结果完全相同,不符合每个分组对应独立结果的预期:

library(dplyr)
calc_cor <- function(df) {
  list(
    Cor1x2 = cor(df$var1, df$var2, method = "spearman", use = "pairwise.complete.obs"),
    Cor1x3 = cor(df$var1, df$var3, method = "spearman", use = "pairwise.complete.obs"),
    Cor1x4 = cor(df$var1, df$var4, method = "spearman", use = "pairwise.complete.obs"),
    Cor1x5 = cor(df$var1, df$var5, method = "spearman", use = "pairwise.complete.obs"),
    Cor2x3 = cor(df$var2, df$var3, method = "spearman", use = "pairwise.complete.obs"),
    Cor2x4 = cor(df$var2, df$var4, method = "spearman", use = "pairwise.complete.obs"),
    Cor2x5 = cor(df$var2, df$var5, method = "spearman", use = "pairwise.complete.obs"),
    Cor3x4 = cor(df$var3, df$var4, method = "spearman", use = "pairwise.complete.obs"),
    Cor3x5 = cor(df$var3, df$var5, method = "spearman", use = "pairwise.complete.obs"),
    Cor4x5 = cor(df$var4, df$var5, method = "spearman", use = "pairwise.complete.obs")
  )
}

result <- df %>%
  group_by(group, meeting, session_phase) %>%
  summarise(correlations = list(calc_cor(.))) %>%
  unnest_wider(correlations) %>%
  as.data.frame()

输入数据中var4和var5存在大量NA值,分组变量group取值1-55,meeting取值1-6,session_phase取值1-2。

原因分析

  1. 函数返回格式与summarise适配问题:原代码中calc_cor返回列表,再用list()包裹后传入summarise,导致unnest_wider未正确解析分组后的独立结果,反而复用了全局数据的计算值。
  2. 分组观测数不足:若部分分组的有效观测数≤1,cor函数会返回固定值(如1或NA),导致多个分组结果看似一致。

修正方案

方案1:调整函数返回格式与调用逻辑

将calc_cor改为返回tibble,直接在summarise中展开结果,无需额外嵌套与反嵌套操作:

library(dplyr)

calc_cor <- function(df) {
  tibble(
    Cor1x2 = cor(df$var1, df$var2, method = "spearman", use = "pairwise.complete.obs"),
    Cor1x3 = cor(df$var1, df$var3, method = "spearman", use = "pairwise.complete.obs"),
    Cor1x4 = cor(df$var1, df$var4, method = "spearman", use = "pairwise.complete.obs"),
    Cor1x5 = cor(df$var1, df$var5, method = "spearman", use = "pairwise.complete.obs"),
    Cor2x3 = cor(df$var2, df$var3, method = "spearman", use = "pairwise.complete.obs"),
    Cor2x4 = cor(df$var2, df$var4, method = "spearman", use = "pairwise.complete.obs"),
    Cor2x5 = cor(df$var2, df$var5, method = "spearman", use = "pairwise.complete.obs"),
    Cor3x4 = cor(df$var3, df$var4, method = "spearman", use = "pairwise.complete.obs"),
    Cor3x5 = cor(df$var3, df$var5, method = "spearman", use = "pairwise.complete.obs"),
    Cor4x5 = cor(df$var4, df$var5, method = "spearman", use = "pairwise.complete.obs")
  )
}

# 计算分组相关性,同时过滤观测数不足2的分组
result <- df %>%
  group_by(group, meeting, session_phase) %>%
  filter(n() >= 2) %>% # 至少2个观测才能计算有效相关性
  summarise(calc_cor(.), .groups = "drop") %>%
  as.data.frame()

方案2:简化相关性计算逻辑(可选)

通过矩阵计算批量生成两两相关性,代码更简洁且不易出错:

library(dplyr)
library(tidyr)

calc_cor_matrix <- function(df) {
  # 提取var1-var5列,计算Spearman相关矩阵
  cor_mat <- cor(df %>% select(starts_with("var")), 
                 method = "spearman", 
                 use = "pairwise.complete.obs")
  # 提取上三角的两两相关值,排除对角线
  cor_pairs <- cor_mat[upper.tri(cor_mat)]
  # 生成对应的列名
  pair_names <- combn(paste0("var", 1:5), 2, function(x) paste0("Cor", x[1], "x", x[2]))
  # 返回tibble格式结果
  as_tibble(setNames(list(cor_pairs), pair_names))
}

result <- df %>%
  group_by(group, meeting, session_phase) %>%
  filter(n() >= 2) %>%
  summarise(calc_cor_matrix(.), .groups = "drop") %>%
  as.data.frame()

验证步骤

  1. 先检查每个分组的观测数,确认是否存在大量样本量不足的情况:
df %>%
  group_by(group, meeting, session_phase) %>%
  tally(name = "obs_count") %>%
  arrange(obs_count)
  1. 运行修正后的代码,随机抽取几个分组的结果,验证相关性是否存在差异。

内容的提问来源于stack exchange,提问作者Jan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.26 05:01:04