You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在ggplot中结合Deposit和CL_Phase绘制盐度直方图?

问题描述

需要用ggplot绘制以**Salinity(盐度)**为x轴的直方图,同时按Deposit和CL_Phase两个变量分类,不想仅靠轮廓颜色区分(辨识度低),但以下代码无法运行:

Salinity <- df %>% ggplot(aes(x = Salinity, fill = (CL_Phase,Deposit))) +
  geom_histogram(color = "black", binwidth = 10)

样本数据集

structure(list(Deposit = c("KA", "KA", "KA", "KA", "KA", "KA", 
"KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "KA", "LS", 
"LS", "LS", "LS", "LS", "LS", "TF", "LS", "LS", "LS", "LS", "LS", 
"TF", "TF", "TF", "LS", "TF", "TF", "TF", "TF", "TF", "TF", "TF", 
"TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", 
"TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", 
"TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", 
"TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", "TF", 
"TF", "TF", "TF"), Salinity = c(1.905259367, 4.494745431, 7.864625, 7.864625, 
8.945789384, 9.855516125, 10.977317768, 12.047547359, 12.163384128, 
12.278617625, 12.845853, 13.937, 14.043035657, 14.461747125, 
17.522345621, 18.717681707, 29.7852583812417, 29.8328258147691, 
29.8368035674166, 30.0793702312925, 30.3339066496739, 30.3806645295019, 
30.6976763213867, 31.1790550169573, 31.2257750014753, 31.3435723173383, 
32.8063329077369, 33.2203487697482, 34.0674202932429, 34.4748269145405, 
34.6361603852619, 34.7049665592227, 34.8689059529115, 34.9515673252259, 
35.0603621635837, 35.0603621635837, 35.0925119138281, 35.1247305837681, 
35.1893750405706, 35.2153102736838, 35.2803423820001, 35.3194949209651, 
35.4573184402101, 35.5631542872796, 35.5697928649395, 35.6163415886054, 
35.7165533901774, 35.7366719228576, 35.8039174991621, 35.9120999790248, 
35.9392592522252, 35.9732724029392, 36.0483522801742, 36.1237779532062, 
36.17885070046, 36.1926475811928, 36.1995503293744, 36.2618044191618, 
36.3103854573673, 36.3312491422852, 36.4009824223941, 36.4429611760086, 
36.4569772567865, 36.5201931269833, 36.8972506639992, 36.9044439812344, 
36.9837645609842, 37.2678697934722, 37.6087773415271, 37.6162586135865, 
38.0400125033386, 38.3407352586274, 39.0042541315711, 39.0766355908795, 
39.189728130044, 39.706196639173, 40.4741922831436, 40.9335612547648, 
41.0827995678457, 41.1180484407973, 42.2730680102387, 44.3120939549952, 
45.299440918469), CL_Phase = c("1", "3", "3", "1", "2", "2", 
"3", "1", "3", "1", "3", "1", "2", "2", "2", "2", "1", "1", "1", 
"1", "1", "1", "3", "1", "1", "1", "1", "1", "1", "3", "3", "1", 
"1", "3", "1", "3", "1", "3", "1", "1", "1", "1", "3", "3", "1", 
"3", "1", "1", "1", "1", "3", "1", "3", "1", "1", "3", "3", "1", 
"3", "1", "1", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", 
"3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3", "3")), row.names = c(NA, 
-83L), class = c("tbl_df", "tbl", "data.frame"))
解决方案

原代码报错是因为fill = (CL_Phase,Deposit)的写法不符合ggplot语法,无法识别双变量组合。以下是几种高辨识度的可视化方案:

1. 分面展示(最清晰的分类方式)

将两个变量分别作为行/列分面,每个子图对应一组Deposit+CL_Phase的组合,完全避免混淆:

library(ggplot2)
library(dplyr)

df %>%
  ggplot(aes(x = Salinity, fill = interaction(CL_Phase, Deposit))) +
  geom_histogram(color = "black", binwidth = 10) +
  facet_grid(Deposit ~ CL_Phase) +
  labs(fill = "CL_Phase + Deposit", x = "盐度", y = "频数") +
  theme_bw()

也可以用facet_wrap按组合分面,适合组合数量较多的场景:

df %>%
  mutate(group = paste(CL_Phase, Deposit, sep = "-")) %>%
  ggplot(aes(x = Salinity, fill = group)) +
  geom_histogram(color = "black", binwidth = 10) +
  facet_wrap(~group) +
  labs(fill = "分组", x = "盐度", y = "频数") +
  theme_bw()

2. 堆叠直方图

将双变量组合作为填充色,在同一图中堆叠展示,适合对比整体分布占比:

df %>%
  ggplot(aes(x = Salinity, fill = interaction(CL_Phase, Deposit))) +
  geom_histogram(color = "black", binwidth = 10, position = "stack") +
  labs(fill = "CL_Phase + Deposit", x = "盐度", y = "频数") +
  theme_bw()

3. 并排分组直方图

每个盐度区间的不同组合并排展示,适合直接对比组间频数差异:

df %>%
  ggplot(aes(x = Salinity, fill = interaction(CL_Phase, Deposit))) +
  geom_histogram(color = "black", binwidth = 10, position = "dodge") +
  labs(fill = "CL_Phase + Deposit", x = "盐度", y = "频数") +
  theme_bw()

关键说明

  • 用interaction(CL_Phase, Deposit)或paste(CL_Phase, Deposit)生成双变量组合,让ggplot能识别填充色的分组依据。
  • 分面方案辨识度最高,适合需要清晰区分每个组合的场景;堆叠/并排方案适合对比整体或组间分布特征。

内容的提问来源于stack exchange,提问作者jlima.geol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 13:54:21