You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多条件箱线图绘制问题:如何按分组正确可视化实验数据

解决多条件箱线图分组逻辑问题

你的核心问题是原始数据处理没有拆分出明确的分组变量,导致ggplot只能把variable列里的Start_0.05这类复合标签当成独立x轴类别,无法体现主条件(Condition1-4)、时间(Start/24h)、浓度(0.05/0.1/0.5)的层级关系。

解决步骤:

1. 拆分出独立分组变量

先从原始数据里提取3个关键分组维度:

  • 主条件(Condition1-4):从X列的ConditionX_rN格式中提取
  • 时间(Start/24h):从variable列的Start_xxx/X24h_xxx格式中提取
  • 浓度(0.05/0.1/0.5):从variable列的后缀提取

2. 完整代码实现

# 加载依赖包(tidyverse包含reshape2、dplyr、tidyr、ggplot2)
library(tidyverse)

# 原始数据
df <- structure(list(X = c("Condition1_r1", "Condition1_r2", "Condition1_r3", 
"Condition2_r1", "Condition2_r2", "Condition2_r3", "Condition3_r1", 
"Condition3_r2", "Condition3_r3", "Condition4_r1", "Condition4_r2", 
"Condition4_r3"), Start_0.05 = c(1985690L, 1648182L, 1753785L, 
1503562L, 1766865L, 1668165L, 1667551L, 1696531L, 1641864L, 1948220L, 
1746022L, 1684197L), X24h_0.05 = c(1806206L, 1808040L, 1891484L, 
1625457L, 1941141L, 1861851L, 1779449L, 1868057L, 1826050L, 1937257L, 
1904913L, 1860859L), Start_0.1 = c(1730114L, 1725849L, 1789878L, 
1740597L, 1701077L, 1713350L, 1707332L, 1783543L, 1797749L, 1610270L, 
2044620L, 1841091L), X24h_0.1 = c(1834778L, 1869877L, 1937330L, 
1888427L, 1886134L, 1902281L, 1835693L, 1937490L, 1933643L, 1716316L, 
1953273L, 1942712L), Start_0.5 = c(1755666L, 1668527L, 1695593L, 
1751354L, 1720210L, 1657771L, 2257244L, 1686353L, 1645991L, 1782086L, 
1785223L, 1793709L), X24h_0.5 = c(1865112L, 1830708L, 1863313L, 
1901987L, 1901676L, 1857649L, 2012898L, 1849991L, 1827155L, 1858165L, 
1930146L, 1958237L)), class = "data.frame", row.names = c(NA, 
-12L))

# 数据预处理:拆分分组变量
df_processed <- df %>%
  # 从X列提取主条件(Condition1-4)
  mutate(Condition = str_extract(X, "Condition\\d+")) %>%
  # 宽表转长表,保留Condition列
  pivot_longer(cols = -c(X, Condition), names_to = "variable", values_to = "value") %>%
  # 拆分variable列为时间和浓度
  separate(variable, into = c("Time", "Concentration"), sep = "_") %>%
  # 把X24h替换成24h,统一时间格式
  mutate(Time = str_replace(Time, "X24h", "24h"))

# 绘制分层箱线图:按主条件分面,x轴为时间,颜色区分浓度
ggplot(df_processed, aes(x = Time, y = value, fill = Concentration)) +
  geom_boxplot(position = position_dodge(width = 0.8)) +
  # 添加散点展示重复样本
  geom_jitter(position = position_dodge(width = 0.8), size = 2, alpha = 0.6) +
  # 按主条件分面,每行2个
  facet_wrap(~Condition, nrow = 2) +
  labs(title = "多条件分组箱线图", x = "时间", y = "数值", fill = "浓度") +
  theme_minimal()

代码说明:

  • str_extract:从X列提取主条件标签,避免重复样本干扰分组
  • pivot_longer:替代melt(tidyverse推荐函数),将宽表转为长表,同时保留主条件信息
  • separate:把复合的variable标签拆成独立的Time和Concentration变量
  • 绘图时用分面展示主条件,x轴展示时间,颜色区分浓度,完美匹配你的分组逻辑;添加geom_jitter可以直观看到每个重复样本的分布

内容的提问来源于stack exchange,提问作者nik

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 23:21:11