如何在ggplot中结合Deposit和CL_Phase绘制盐度直方图?
问题描述
需要用ggplot绘制以**Salinity(盐度)**为x轴的直方图,同时按Deposit和CL_Phase两个变量分类,不想仅靠轮廓颜色区分(辨识度低),但以下代码无法运行:
Salinity <- df %>% ggplot(aes(x = Salinity, fill = (CL_Phase,Deposit))) + geom_histogram(color = "black", binwidth = 10)
样本数据集
structure(list(Deposit = c("KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "LS", "LS", "LS", "LS", "LS", "LS", "TF", "LS", "LS", "LS", "LS", "LS", "TF", "TF", "TF", "LS", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF"), Salinity = c(1.905259367, 4.494745431, 7.864625, 7.864625, 8.945789384, 9.855516125, 10.977317768, 12.047547359, 12.163384128, 12.278617625, 12.845853, 13.937, 14.043035657, 14.461747125, 17.522345621, 18.717681707, 29.7852583812417, 29.8328258147691, 29.8368035674166, 30.0793702312925, 30.3339066496739, 30.3806645295019, 30.6976763213867, 31.1790550169573, 31.2257750014753, 31.3435723173383, 32.8063329077369, 33.2203487697482, 34.0674202932429, 34.4748269145405, 34.6361603852619, 34.7049665592227, 34.8689059529115, 34.9515673252259, 35.0603621635837, 35.0603621635837, 35.0925119138281, 35.1247305837681, 35.1893750405706, 35.2153102736838, 35.2803423820001, 35.3194949209651, 35.4573184402101, 35.5631542872796, 35.5697928649395, 35.6163415886054, 35.7165533901774, 35.7366719228576, 35.8039174991621, 35.9120999790248, 35.9392592522252, 35.9732724029392, 36.0483522801742, 36.1237779532062, 36.17885070046, 36.1926475811928, 36.1995503293744, 36.2618044191618, 36.3103854573673, 36.3312491422852, 36.4009824223941, 36.4429611760086, 36.4569772567865, 36.5201931269833, 36.8972506639992, 36.9044439812344, 36.9837645609842, 37.2678697934722, 37.6087773415271, 37.6162586135865, 38.0400125033386, 38.3407352586274, 39.0042541315711, 39.0766355908795, 39.189728130044, 39.706196639173, 40.4741922831436, 40.9335612547648, 41.0827995678457, 41.1180484407973, 42.2730680102387, 44.3120939549952, 45.299440918469), CL_Phase = c("1", "3", "3", "1", "2", "2", "3", "1", "3", "1", "3", "1", "2", "2", "2", "2", "1", "1", "1", "1", "1", "1", "3", "1", "1", "1", "1", "1", "1", "3", "3", "1", "1", "3", "1", "3", "1", "3", "1", "1", "1", "1", "3", "3", "1", "3", "1", "1", "1", "1", "3", "1", "3", "1", "1", "3", "3", "1", "3", "1", "1", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3")), row.names = c(NA, -83L), class = c("tbl_df", "tbl", "data.frame"))
解决方案
原代码报错是因为fill = (CL_Phase,Deposit)的写法不符合ggplot语法,无法识别双变量组合。以下是几种高辨识度的可视化方案:
1. 分面展示(最清晰的分类方式)
将两个变量分别作为行/列分面,每个子图对应一组Deposit+CL_Phase的组合,完全避免混淆:
library(ggplot2) library(dplyr) df %>% ggplot(aes(x = Salinity, fill = interaction(CL_Phase, Deposit))) + geom_histogram(color = "black", binwidth = 10) + facet_grid(Deposit ~ CL_Phase) + labs(fill = "CL_Phase + Deposit", x = "盐度", y = "频数") + theme_bw()
也可以用facet_wrap按组合分面,适合组合数量较多的场景:
df %>% mutate(group = paste(CL_Phase, Deposit, sep = "-")) %>% ggplot(aes(x = Salinity, fill = group)) + geom_histogram(color = "black", binwidth = 10) + facet_wrap(~group) + labs(fill = "分组", x = "盐度", y = "频数") + theme_bw()
2. 堆叠直方图
将双变量组合作为填充色,在同一图中堆叠展示,适合对比整体分布占比:
df %>% ggplot(aes(x = Salinity, fill = interaction(CL_Phase, Deposit))) + geom_histogram(color = "black", binwidth = 10, position = "stack") + labs(fill = "CL_Phase + Deposit", x = "盐度", y = "频数") + theme_bw()
3. 并排分组直方图
每个盐度区间的不同组合并排展示,适合直接对比组间频数差异:
df %>% ggplot(aes(x = Salinity, fill = interaction(CL_Phase, Deposit))) + geom_histogram(color = "black", binwidth = 10, position = "dodge") + labs(fill = "CL_Phase + Deposit", x = "盐度", y = "频数") + theme_bw()
关键说明
- 用
interaction(CL_Phase, Deposit)或paste(CL_Phase, Deposit)生成双变量组合,让ggplot能识别填充色的分组依据。 - 分面方案辨识度最高,适合需要清晰区分每个组合的场景;堆叠/并排方案适合对比整体或组间分布特征。
内容的提问来源于stack exchange,提问作者jlima.geol
相关产品推荐
相关产品推荐

