堆叠条形图:如何添加样本量数值与统计检验结果?
问题1:在条形顶部添加样本量
要给每个颜色的堆叠条形顶部标注样本量,需先单独统计每组样本量,再用geom_text()将标签定位到条形顶部:
# 加载依赖包 library(ggplot2) library(dplyr) library(ggsci) # 计算各颜色组的样本量 sample_sizes <- my_data %>% count(flower_color, name = "sample_size") # 绘制带样本量的堆叠条形图 table(my_data$flower_color, my_data$flower_length) %>% as.data.frame() %>% filter(Freq > 0) %>% ggplot(aes(x = Var1, y = Freq, fill = Var2)) + geom_bar(position = "fill", stat = "identity") + # 添加样本量标签,position_fill(vjust=1.05)确保标签在条形上方 geom_text(data = sample_sizes, aes(x = flower_color, y = 1, label = sample_size, fill = NULL), position = position_fill(vjust = 1.05), size = 4) + scale_y_continuous(labels = scales::percent_format()) + theme(panel.grid.major = element_blank(), panel.grid.minor = element_blank(), panel.background = element_blank(), axis.line = element_line(colour = "black")) + scale_fill_jco() # 替代原代码中的set_palette,需加载ggsci包
问题2:统计检验与显著性添加
你的数据是分类变量的列联表(花颜色与花长度的交叉分布),适合用**卡方检验(Chi-squared test)**验证组间分布差异;若单元格期望频数小于5则改用Fisher精确检验,你的数据集样本量足够,卡方检验适用。
方法1:添加全局卡方检验p值
用ggpubr包的stat_compare_means()直接在图中添加全局检验结果:
library(ggpubr) table(my_data$flower_color, my_data$flower_length) %>% as.data.frame() %>% filter(Freq > 0) %>% ggplot(aes(x = Var1, y = Freq, fill = Var2)) + geom_bar(position = "fill", stat = "identity") + geom_text(data = sample_sizes, aes(x = flower_color, y = 1, label = sample_size, fill = NULL), position = position_fill(vjust = 1.05), size = 4) + scale_y_continuous(labels = scales::percent_format()) + theme(panel.grid.major = element_blank(), panel.grid.minor = element_blank(), panel.background = element_blank(), axis.line = element_line(colour = "black")) + scale_fill_jco() + # 添加全局卡方检验p值,label.y控制垂直位置 stat_compare_means(method = "chisq", label = "p.format", label.y = 1.1)
方法2:添加组间两两比较的显著性标记
如果需要对比特定组(如蓝色与红色、蓝色与粉色)的分布差异,用pairwise.prop.test做事后检验,再用ggsignif的geom_signif()添加标记:
library(ggsignif) # 执行两两比例检验,用bonferroni调整p值 pairwise_result <- pairwise.prop.test(table(my_data$flower_color, my_data$flower_length), p.adjust.method = "bonferroni") # 整理比较结果 compare_df <- data.frame( group1 = c("blue", "blue"), group2 = c("red", "pink"), p_val = pairwise_result$p.value[lower.tri(pairwise_result$p.value)], y_pos = c(1.05, 1.1) ) %>% mutate(signif = case_when( p_val < 0.001 ~ "***", p_val < 0.01 ~ "**", p_val < 0.05 ~ "*", TRUE ~ "ns" )) # 绘图添加显著性标记 table(my_data$flower_color, my_data$flower_length) %>% as.data.frame() %>% filter(Freq > 0) %>% ggplot(aes(x = Var1, y = Freq, fill = Var2)) + geom_bar(position = "fill", stat = "identity") + geom_text(data = sample_sizes, aes(x = flower_color, y = 1, label = sample_size, fill = NULL), position = position_fill(vjust = 1.05), size = 4) + scale_y_continuous(labels = scales::percent_format()) + theme(panel.grid.major = element_blank(), panel.grid.minor = element_blank(), panel.background = element_blank(), axis.line = element_line(colour = "black")) + scale_fill_jco() + geom_signif( data = compare_df, aes(xmin = group1, xmax = group2, annotations = signif, y_position = y_pos), tip_length = 0.01, manual = TRUE )
内容的提问来源于stack exchange,提问作者Edoardo Scali
相关产品推荐
相关产品推荐

