R语言geom_bar绘制占比求助:以总观测数为100%
解决你的R语言ggplot绘图问题
没问题,我来帮你搞定这两个问题——让百分比以总观测数为基准,以及正确添加频率标签!
问题1:让百分比以总观测数为基准
你之前的代码用了subset(my.data, PHENO == 2),这会让ggplot只处理PHENO=2的观测,所以..prop..计算的是PHENO=2内部各bin的比例,而不是相对于所有观测(PHENO=1+2)的比例。要修正这个,我们需要基于完整数据集计算每个bin中PHENO=2的数量占总观测数的比例。
有两种常用方法:
方法1:直接在ggplot中计算比例
利用after_stat()函数在绘图时实时计算比例:
library(ggplot2) library(scales) ggplot(data = my.data, aes(x = as.factor(bins))) + # 绘制柱状图,y为PHENO=2的数量/总观测数 geom_bar(aes(y = after_stat(count[PHENO == 2])/nrow(my.data)), stat = "count", subset = .(PHENO == 2)) + # 设置y轴为百分比格式 scale_y_continuous(labels = percent_format(), limits = c(0, 0.15)) + # 保留你的参考线和注释 geom_hline(yintercept = 0.05, linetype="dashed", color = 'blue', size = 1) + annotate(geom = "text", label = 'Prevalence 5%', x = 1.5, y = 0.05, vjust = -1, col = 'blue')
方法2:先预处理数据(更直观)
先计算好每个bin的比例,再绘图,这样更容易验证数值是否正确:
library(ggplot2) library(scales) library(dplyr) # 预处理:计算每个bin中PHENO=2的百分比(基于总观测数) summary_data <- my.data %>% group_by(bins) %>% summarise( pheno2_count = sum(PHENO == 2), total_obs = nrow(my.data), pheno2_pct = pheno2_count / total_obs ) # 用预处理后的数据绘图 ggplot(summary_data, aes(x = as.factor(bins), y = pheno2_pct)) + geom_col(fill = "steelblue") + scale_y_continuous(labels = percent_format(), limits = c(0, 0.15)) + geom_hline(yintercept = 0.05, linetype="dashed", color = 'blue', size = 1) + annotate(geom = "text", label = 'Prevalence 5%', x = 1.5, y = 0.05, vjust = -1, col = 'blue')
问题2:添加正确的频率标签
你之前的标签用了as.factor(bins),所以显示的是bin编号,而不是频率。我们需要把标签改成计算好的百分比,同样分两种方法:
对应方法1的标签代码
在原ggplot代码中添加geom_text,用after_stat()计算标签内容:
# 接方法1的代码,添加这部分 geom_text(aes(y = after_stat(count[PHENO == 2])/nrow(my.data)), stat = "count", subset = .(PHENO == 2), label = scales::percent(after_stat(count[PHENO == 2])/nrow(my.data)), vjust = -0.25)
对应方法2的标签代码
用预处理好的summary_data,直接调用计算好的百分比:
# 接方法2的代码,添加这部分 geom_text(aes(label = percent(pheno2_pct)), vjust = -0.25)
完整示例代码(预处理版,推荐)
library(ggplot2) library(scales) library(dplyr) # 模拟你的数据(方便测试) set.seed(123) my.data <- data.frame( PHENO = factor(sample(c(1,2), 1000, replace = TRUE, prob = c(0.95, 0.05))), bins = sample(1:10, 1000, replace = TRUE) ) # 预处理数据 summary_data <- my.data %>% group_by(bins) %>% summarise( pheno2_count = sum(PHENO == 2), total_obs = nrow(my.data), pheno2_pct = pheno2_count / total_obs ) # 最终绘图 ggplot(summary_data, aes(x = as.factor(bins), y = pheno2_pct)) + geom_col(fill = "steelblue") + scale_y_continuous(labels = percent_format(), limits = c(0, 0.15)) + geom_hline(yintercept = 0.05, linetype="dashed", color = 'blue', size = 1) + annotate(geom = "text", label = 'Prevalence 5%', x = 1.5, y = 0.05, vjust = -1, col = 'blue') + geom_text(aes(label = percent(pheno2_pct)), vjust = -0.25) + labs(x = "Bins", y = "Percentage of Total Observations (PHENO=2)")
这样绘制出来的图表,y轴的百分比就是基于所有观测数的,每个柱子上方也会显示对应的百分比标签啦!
内容的提问来源于stack exchange,提问作者tibetish
相关产品推荐
相关产品推荐

