含重复实验数据的柱状图+散点图绘制问题求助
问题
数据如下:
| group | condition | value | Source |
|---|---|---|---|
| A | Health | 1.0000000 | Simulated |
| B | Health | 1.0000000 | Simulated |
| C | Health | 1.0000000 | Simulated |
| A | Disease | 0.4589925 | Simulated |
| B | Disease | 1.4905422 | Simulated |
| C | Disease | 1.3422035 | Simulated |
| A | Health | 1.0000000 | Experimental |
| B | Health | 1.0000000 | Experimental |
| C | Health | 1.0000000 | Experimental |
| A | Disease | 0.5158371 | Experimental |
| A | Disease | 0.7559055 | Experimental |
| B | Disease | 1.4153005 | Experimental |
| B | Disease | 1.4000000 | Experimental |
| C | Disease | 1.3300971 | Experimental |
| C | Disease | 1.0000000 | Experimental |
需求:绘制组合图,X轴为Group,Y轴为value;每个组包含两个柱状图(对应Health和Disease条件,使用Simulated数据),同时叠加Experimental数据的散点图。
此前当每个条件仅有一条Experimental数据时,使用以下代码实现:
pivoted<-data %>% pivot_wider(names_from =Source) ggplot(pivoted, aes(x=group,fill=condition)) + geom_col(aes(y=Simulated), position=position_dodge()) + labs(x="Group", y="Value") + geom_point(aes(y=Experimental), position = position_dodge(width=0.9), size=3)
但现在Experimental数据存在重复值,上述方法无法正常工作,寻求解决办法。
解决方案
你之前的代码失效是因为pivot_wider遇到重复的Experimental记录时,会自动生成多列(比如Experimental_1、Experimental_2),没法直接用y=Experimental映射。不需要转宽表,直接基于原始长表分别处理两类数据就行:
library(ggplot2) library(dplyr) ggplot() + # 绘制Simulated数据的柱状图 geom_col(data = data %>% filter(Source == "Simulated"), aes(x = group, y = value, fill = condition), position = position_dodge(width = 0.9)) + # 绘制Experimental数据的散点,用position_jitterdodge实现对齐+防重叠 geom_point(data = data %>% filter(Source == "Experimental"), aes(x = group, y = value, color = condition), position = position_jitterdodge(jitter.width = 0.1, dodge.width = 0.9), size = 3) + labs(x = "Group", y = "Value") + theme_minimal()
要点说明:
- 给每个geom单独指定对应的数据子集,不用折腾宽表转换
position_jitterdodge既能让散点和同组同条件的柱子精准对齐,又能让重叠的散点稍微错开,方便看清所有实验数据- 给散点设置和柱状图不同的颜色,能增强两类数据的区分度,视觉效果更清晰
内容的提问来源于stack exchange,提问作者AFP
相关产品推荐
相关产品推荐

