为geom_col()添加百分比标签:Pima Indians数据集可视化问题排查
修正Pima Indians数据集可视化问题
问题说明
当前代码存在两个问题:
- 百分比计算逻辑错误:原代码将特征总和除以数据集总行数,这不是合理的占比统计,正确逻辑应为每个糖尿病类别(Yes/No)的特征总和,占该特征所有样本总和的百分比
- 百分比文本颜色未生效:原代码将
color="white"放在aes()内部,导致颜色被当作变量映射而非固定值
修正后代码
library(MASS) library(dplyr) library(tidyr) library(ggplot2) Pima.te %>% mutate(across(-type, as.numeric)) %>% pivot_longer(-type, names_to = "var", values_to = "value") %>% summarise(value = sum(value), .by = c(type, var)) %>% # 按特征分组计算总合,再计算百分比 group_by(var) %>% mutate(total = sum(value), percentage = round(value / total * 100, 2)) %>% ungroup() %>% ggplot(aes(x = type, y = value)) + geom_col() + geom_text(aes(label = value), vjust = -.2) + # 固定设置文本颜色为白色 geom_text(aes(label = paste0(percentage, "%")), vjust = 3, color = "white") + scale_y_continuous(expand = c(0, 0, .2, 0)) + facet_wrap(~var, scales = "free") + labs(title = "糖尿病类别与各特征总和对比(含百分比)")
修正点解释
- 百分比计算修正:
- 先按
var(特征)分组,计算每个特征的总总和total - 用
value / total * 100得到每个类别在该特征下的占比,转为百分比格式后保留两位小数
- 先按
- 文本颜色修正:
- 将
color="white"从aes()中移出,作为geom_text()的直接参数,确保文本固定显示为白色
- 将
内容的提问来源于stack exchange,提问作者Russ Conte
相关产品推荐
相关产品推荐

