使用ggplot2绘制非数值分类数据的可视化方案咨询
适配需求的ggplot2可视化实现方案
方案1:分组丰度热图
适用于同时展示6个组别、20个分类项的丰度等级全局分布,是科研场景中多分类组间对比的常用方案,可直观呈现不同组的分类丰度规律。
先对丰度等级做有序因子处理,匹配丰度从高到低的逻辑:
library(tidyverse) # 数据预处理 df_long <- df %>% pivot_longer(starts_with("cat_"), names_to = "category") %>% mutate(value = ifelse(value == "-", NA, value), # 定义丰度等级顺序 value = factor(value, levels = c("D", "Mj", "Mn", "Tr"))) %>% drop_na() %>% # 统计每组+每个分类+每个丰度等级的出现频次 count(group, category, value, name = "freq") # 绘制热图 ggplot(df_long, aes(x = category, y = group, fill = value)) + geom_tile(color = "white", size = 0.3) + # 可叠加频次文本 geom_text(aes(label = freq), color = "black", size = 3) + scale_fill_brewer(palette = "RdYlBu", direction = -1, na.value = "grey90", labels = c("D: >50%", "Mj: 5%-50%", "Mn: 1%-5%", "Tr: <1%")) + labs(x = "分类项", y = "组别", fill = "丰度等级") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
方案2:堆叠百分比条形图
适用于对比不同组别下,同一分类项的丰度等级占比差异,可直接呈现组间分布的异质性。
ggplot(df_long, aes(x = category, y = freq, fill = value)) + geom_col(position = "fill") + facet_wrap(~ group, ncol = 2) + scale_y_continuous(labels = scales::percent_format()) + scale_fill_brewer(palette = "RdYlBu", direction = -1, labels = c("D: >50%", "Mj: 5%-50%", "Mn: 1%-5%", "Tr: <1%")) + labs(x = "分类项", y = "占比", fill = "丰度等级") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
方案3:分面抖动散点图
适用于需要保留单样本层面的原始观测信息的场景,适配组内样本量不固定的情况,不会因汇总丢失原始数据特征。
# 保留样本维度的预处理 df_sample_long <- df %>% pivot_longer(starts_with("cat_"), names_to = "category") %>% mutate(value = ifelse(value == "-", NA, value), value = factor(value, levels = c("D", "Mj", "Mn", "Tr"))) %>% drop_na() ggplot(df_sample_long, aes(x = value, y = category, color = value)) + geom_jitter(width = 0.2, height = 0.2, alpha = 0.7, size = 2) + facet_wrap(~ group, ncol = 2) + scale_color_brewer(palette = "RdYlBu", direction = -1, labels = c("D: >50%", "Mj: 5%-50%", "Mn: 1%-5%", "Tr: <1%")) + labs(x = "丰度等级", y = "分类项", color = "丰度等级") + theme_minimal()
内容的提问来源于stack exchange,提问作者aterhorst
相关产品推荐
相关产品推荐

