You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中为多面板脑区体积箱线图添加ANOVA统计比较

问题:在多面板箱线图中添加ANOVA组间统计比较结果

我有一份CSV格式的脑区体积数据,样本结构包含基因型(16pDeletion、16pDuplication,计划补充Control组)和多个脑区体积指标(TCV、cGMV、sGMV、WMV、Ventricles)。我希望绘制多面板箱线图展示各脑区在三种基因型下的体积分布,并添加组间ANOVA统计结果,但现有代码仅能生成箱线图,无法显示统计信息。

现有代码如下:

#Open Libraries 
library(ggplot2)
library(dplyr)
library(tidyr)
library(car)
library(ggpubr)

# Read in the data
data <- read.csv("SynthSegQC.csv")

# Manually specify the columns for the regions of interest
regions_of_interest <- c("TCV", "cGMV", "sGMV", "WMV", "Ventricles" ) 

# Reshape the data to long format
data_long <- pivot_longer(data, cols = all_of(regions_of_interest), names_to = "Region", values_to = "Volume")

# Function to perform ANOVA and return p-values
get_p_value <- function(data, region) {
  model <- aov(Volume ~ Genotype, data = data[data$Region == region,])
  p_value <- summary(model)[[1]]$"Pr(>F)"[1]
  return(p_value)
}

# Add a column with p-values for each region
data_long <- data_long %>%
  group_by(Region) %>%
  mutate(P_Value = get_p_value(data_long, Region))

# Function to format p-values for display
format_p_value <- function(p) {
  ifelse(p < 0.001, "<0.001", sprintf("%.3f", p))
}

# Ensure that the p-values are calculated outside of ggplot2 aesthetics
data_long$Formatted_P_Values <- format_p_value(data_long$P_Value)

# Plotting with facets
ggplot(data_long, aes(x = Genotype, y = Volume, fill = Genotype)) +
  geom_boxplot() +
  geom_jitter(aes(color = Genotype), width = 0.2, size = 1, alpha = 0.5) +
  facet_wrap(~Region, scales = "free") +
  theme_minimal() +
  labs(title = "Brain Region Volumes by Genotype", x = "Genotype", y = "Volume") +
  scale_fill_brewer(palette = "Set1") +
  theme(axis.text.x = element_text(angle = 45, hjust = 1)) +
  geom_text(aes(label = Formatted_P_Values, y = Inf), vjust = -0.5, size = 2.5)

解决方案

核心问题分析

原代码手动计算p值后用geom_text添加时,由于每个分组重复携带p值,导致文本重复渲染,且定位逻辑不够灵活。推荐使用ggpubr包的stat_compare_means函数,它能自动在分面面板中计算并添加ANOVA统计结果,无需手动处理数据。

修改后的完整代码

library(ggplot2)
library(dplyr)
library(tidyr)
library(ggpubr)

# 读取数据(若实际包含Control组,直接使用即可)
data <- read.csv("SynthSegQC.csv")

# 指定目标脑区列
regions_of_interest <- c("TCV", "cGMV", "sGMV", "WMV", "Ventricles") 

# 转换为长格式
data_long <- pivot_longer(data, cols = all_of(regions_of_interest), 
                          names_to = "Region", values_to = "Volume")

# 绘制多面板箱线图并添加ANOVA统计
ggplot(data_long, aes(x = Genotype, y = Volume, fill = Genotype)) +
  geom_boxplot(width = 0.6) +
  geom_jitter(aes(color = Genotype), width = 0.2, size = 1, alpha = 0.5) +
  # 分面展示每个脑区,仅Y轴自由缩放以保持X轴一致性
  facet_wrap(~Region, scales = "free_y") +
  # 自动计算每个分面的ANOVA p值并格式化展示
  stat_compare_means(method = "anova", label = "p.format", 
                     label.y = Inf, vjust = 1.2, size = 3) +
  theme_minimal() +
  labs(title = "Brain Region Volumes by Genotype", 
       x = "Genotype", y = "Volume") +
  scale_fill_brewer(palette = "Set1") +
  scale_color_brewer(palette = "Set1") +
  theme(
    axis.text.x = element_text(angle = 45, hjust = 1),
    plot.title = element_text(hjust = 0.5, size = 14),
    strip.text = element_text(size = 12)
  )

关键修改说明

  1. 移除手动p值计算逻辑:stat_compare_means会自动对每个分面(脑区)执行ANOVA分析并提取p值,无需提前计算和存储,避免重复渲染问题。
  2. 优化分面缩放:将scales = "free"改为scales = "free_y",保持X轴(基因型)统一,仅Y轴(体积)根据脑区范围自适应,提升图表可读性。
  3. 统计文本定位:通过label.y = Inf和vjust = 1.2将p值文本固定在面板顶部,避免与箱线图或散点重叠。
  4. 统一配色:为散点设置与箱线图一致的配色方案,增强视觉协调性。

补充说明

若后续补充Control组,代码无需修改,stat_compare_means会自动纳入三组进行ANOVA分析。若需要展示两两比较结果(如Tukey HSD),可添加stat_compare_means(method = "tukey", label = "p.signif", tip.length = 0.01),但需注意分面空间是否足够容纳标记。

内容的提问来源于stack exchange,提问作者SP_2023

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 10:57:02