You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用cor.test()与sm_statCorr()计算Spearman rho结果不一致的问题咨询

问题:cor.test()与sm_statCorr()的Spearman相关系数结果不一致

在对比偏头痛两种症状的两份调查数据相关性时,发现cor.test()计算的Spearman's rho与smplot2包中sm_statCorr()在图表中展示的结果存在细微差异:

  • cor.test()结果:头痛症状rho=0.5028556,先兆症状rho=0.8174239
  • sm_statCorr()图表显示:头痛症状rho=0.55,先兆症状rho=0.83

以下是可复现代码:

install.packages("pacman")
pacman::p_load(stats, dplyr, ggplot2, smplot2) 

# 导入并转换数据
data <- matrix(c(1, "headache", 1, NA, 1, "aura", 0, 0, 2, "headache", 1, 1, 2, "aura", 0, 0, 
                 3, "headache", -1, -1, 3, "aura", 0, 1, 4, "headache", 1, 2, 4, "aura", -2, -2, 
                 5, "headache", 0, 1, 5, "aura", 1, 1, 6, "headache", 2, 2, 6, "aura", 0, 0, 
                 7, "headache", 0, 0, 7, "aura", 0, 0, 8, "headache", 1, 1, 8, "aura", 0, 0, 
                 9, "headache", 1, 0, 9, "aura", 0, 0, 10, "headache", 1, -1, 10, "aura", 0, 0, 
                 11, "headache", 0, 1, 11, "aura", 0, 0, 12, "headache", 0, 0, 12, "aura", 0, 0), 
               nrow = 24, ncol = 4, byrow = T)
colnames(data) = c("id", "symptom", "survey_1", "survey_2")
data <- as.matrix(data)
data <- as.data.frame(data)
data$id <- as.numeric(data$id)
data$survey_1 <- as.numeric(data$survey_1)
data$survey_2 <- as.numeric(data$survey_2)
data$symptom = factor(data$symptom, levels=c("headache", "aura"))

# 使用sm_statCorr绘制图表
data_n <- data %>% group_by(symptom, survey_1, survey_2) %>% tally() %>% ungroup()
data_n <- data_n[,c("symptom", "survey_1", "survey_2", "n")]

ggplot(data_n, aes(x = survey_1, y = survey_2)) + 
  geom_point(aes(size = n)) +
  sm_statCorr(data = data, aes(x = survey_1, y = survey_2), corr_method = "spearman") +
  facet_wrap(. ~ symptom)

# 使用cor.test计算相关系数
cor.test(data[(data$symptom == "headache"), ]$survey_1, 
         data[(data$symptom == "headache"), ]$survey_2,
         method = "spearman")
cor.test(data[(data$symptom == "aura"), ]$survey_1, 
         data[(data$symptom == "aura"), ]$survey_2,
         method = "spearman")

差异原因

核心问题是数据是否保留重复观测:

  • cor.test()使用的是原始全量数据(剔除NA后),包含所有重复的观测记录,每个记录单独参与计算。
  • sm_statCorr()在分面场景下,实际使用的是ggplot主数据data_n(聚合后的去重数据),仅用每个独特的(x,y)点计算相关系数,忽略了重复观测的频数权重。

以头痛症状为例:

  • 原始数据剔除NA后有11条观测,其中包含多个重复的(x,y)组合(如(0,0)出现2次、(0,1)出现2次)。
  • 聚合后的data_n仅保留8个独特的(x,y)点,sm_statCorr()基于这8个点计算出的Spearman rho约为0.55,与全量数据计算的0.5028存在差异。

解决办法

要让两个工具结果一致,需确保sm_statCorr()使用包含重复观测的全量数据:

方法1:直接用原始数据绘图

去掉聚合步骤,用原始数据绘制散点图(可加抖动避免点重叠),sm_statCorr()会自动基于全量数据计算:

ggplot(data, aes(x = survey_1, y = survey_2)) + 
  geom_jitter(width=0.1, height=0.1) +
  sm_statCorr(corr_method = "spearman") +
  facet_wrap(. ~ symptom)

方法2:扩展聚合数据为全量重复数据

如果坚持使用聚合后的点大小展示频数,可先将data_n扩展为包含重复观测的数据集,再传入sm_statCorr():

# 扩展聚合数据为全量重复观测
data_expanded <- data_n %>% uncount(n)

ggplot(data_n, aes(x = survey_1, y = survey_2)) + 
  geom_point(aes(size = n)) +
  sm_statCorr(data = data_expanded, aes(x = survey_1, y = survey_2), corr_method = "spearman") +
  facet_wrap(. ~ symptom)

内容的提问来源于stack exchange,提问作者Wandering_geek

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 18:45:54