使用cor.test()与sm_statCorr()计算Spearman rho结果不一致的问题咨询
问题:cor.test()与sm_statCorr()的Spearman相关系数结果不一致
在对比偏头痛两种症状的两份调查数据相关性时,发现cor.test()计算的Spearman's rho与smplot2包中sm_statCorr()在图表中展示的结果存在细微差异:
cor.test()结果:头痛症状rho=0.5028556,先兆症状rho=0.8174239sm_statCorr()图表显示:头痛症状rho=0.55,先兆症状rho=0.83
以下是可复现代码:
install.packages("pacman") pacman::p_load(stats, dplyr, ggplot2, smplot2) # 导入并转换数据 data <- matrix(c(1, "headache", 1, NA, 1, "aura", 0, 0, 2, "headache", 1, 1, 2, "aura", 0, 0, 3, "headache", -1, -1, 3, "aura", 0, 1, 4, "headache", 1, 2, 4, "aura", -2, -2, 5, "headache", 0, 1, 5, "aura", 1, 1, 6, "headache", 2, 2, 6, "aura", 0, 0, 7, "headache", 0, 0, 7, "aura", 0, 0, 8, "headache", 1, 1, 8, "aura", 0, 0, 9, "headache", 1, 0, 9, "aura", 0, 0, 10, "headache", 1, -1, 10, "aura", 0, 0, 11, "headache", 0, 1, 11, "aura", 0, 0, 12, "headache", 0, 0, 12, "aura", 0, 0), nrow = 24, ncol = 4, byrow = T) colnames(data) = c("id", "symptom", "survey_1", "survey_2") data <- as.matrix(data) data <- as.data.frame(data) data$id <- as.numeric(data$id) data$survey_1 <- as.numeric(data$survey_1) data$survey_2 <- as.numeric(data$survey_2) data$symptom = factor(data$symptom, levels=c("headache", "aura")) # 使用sm_statCorr绘制图表 data_n <- data %>% group_by(symptom, survey_1, survey_2) %>% tally() %>% ungroup() data_n <- data_n[,c("symptom", "survey_1", "survey_2", "n")] ggplot(data_n, aes(x = survey_1, y = survey_2)) + geom_point(aes(size = n)) + sm_statCorr(data = data, aes(x = survey_1, y = survey_2), corr_method = "spearman") + facet_wrap(. ~ symptom) # 使用cor.test计算相关系数 cor.test(data[(data$symptom == "headache"), ]$survey_1, data[(data$symptom == "headache"), ]$survey_2, method = "spearman") cor.test(data[(data$symptom == "aura"), ]$survey_1, data[(data$symptom == "aura"), ]$survey_2, method = "spearman")
差异原因
核心问题是数据是否保留重复观测:
cor.test()使用的是原始全量数据(剔除NA后),包含所有重复的观测记录,每个记录单独参与计算。sm_statCorr()在分面场景下,实际使用的是ggplot主数据data_n(聚合后的去重数据),仅用每个独特的(x,y)点计算相关系数,忽略了重复观测的频数权重。
以头痛症状为例:
- 原始数据剔除NA后有11条观测,其中包含多个重复的(x,y)组合(如(0,0)出现2次、(0,1)出现2次)。
- 聚合后的
data_n仅保留8个独特的(x,y)点,sm_statCorr()基于这8个点计算出的Spearman rho约为0.55,与全量数据计算的0.5028存在差异。
解决办法
要让两个工具结果一致,需确保sm_statCorr()使用包含重复观测的全量数据:
方法1:直接用原始数据绘图
去掉聚合步骤,用原始数据绘制散点图(可加抖动避免点重叠),sm_statCorr()会自动基于全量数据计算:
ggplot(data, aes(x = survey_1, y = survey_2)) + geom_jitter(width=0.1, height=0.1) + sm_statCorr(corr_method = "spearman") + facet_wrap(. ~ symptom)
方法2:扩展聚合数据为全量重复数据
如果坚持使用聚合后的点大小展示频数,可先将data_n扩展为包含重复观测的数据集,再传入sm_statCorr():
# 扩展聚合数据为全量重复观测 data_expanded <- data_n %>% uncount(n) ggplot(data_n, aes(x = survey_1, y = survey_2)) + geom_point(aes(size = n)) + sm_statCorr(data = data_expanded, aes(x = survey_1, y = survey_2), corr_method = "spearman") + facet_wrap(. ~ symptom)
内容的提问来源于stack exchange,提问作者Wandering_geek
相关产品推荐
相关产品推荐

