如何让ggplot2次Y轴在自定义累计值位置设置刻度断点?
如何在ggplot2中实现帕累托图的次Y轴累计值刻度断点?
先来看@camille最初用ggplot2制作的帕累托图代码,效果已经相当不错了:
library(tidyverse) d <- tribble( ~ category, ~defect, "price", 80, "schedule", 27, "supplier", 66, "contact", 94, "item", 33 ) %>% arrange(desc(defect)) %>% mutate( cumsum = cumsum(defect), freq = round(defect / sum(defect), 3), cum_freq = cumsum(freq) ) %>% mutate(category = as.factor(category) %>% fct_reorder(defect)) brks <- unique(d$cumsum) ggplot(d, aes(x = fct_rev(category))) + geom_col(aes(y = defect)) + geom_point(aes(y = cumsum)) + geom_line(aes(y = cumsum, group = 1)) + scale_y_continuous(sec.axis = sec_axis(~. / max(d$cumsum), labels = scales::percent), breaks = brks)
对应的效果如下:
这个图的表现已经近乎完美,但咱们还想优化一点:让次Y轴的刻度精准落在每个累计值的位置——就像下面这段Base R代码实现的效果:
## Creating the d tribble library(tidyverse) d <- tribble( ~ category, ~defect, "price", 80, "schedule", 27, "supplier", 66, "contact", 94, "item", 33 ) ## Creating new columns d <- arrange(d, desc(defect)) %>% mutate( cumsum = cumsum(defect), freq = round(defect / sum(defect), 3), cum_freq = cumsum(freq) ) ## Saving Parameters def_par <- par() ## New margins par(mar=c(5,5,4,5)) ## bar plot, pc will hold x values for bars pc = barplot(d$defect, width = 1, space = 0.2, border = NA, axes = F, ylim = c(0, 1.05 * max(d$cumsum, na.rm = T)), ylab = "Cummulative Counts" , cex.names = 0.7, names.arg = d$category, main = "Pareto Chart (version 1)") ## Cumulative counts line lines(pc, d$cumsum, type = "b", cex = 0.7, pch = 19, col="cyan4") ## Framing plot box(col = "grey62") ## adding axes axis(side = 2, at = c(0, d$cumsum), las = 1, col.axis = "grey62", col = "grey62", cex.axis = 0.8) axis(side = 4, at = c(0, d$cumsum), labels = paste(c(0, round(d$cum_freq * 100)) ,"%",sep=""), las = 1, col.axis = "cyan4", col = "cyan4", cex.axis = 0.8) ## restoring default paramenter par(def_par)
对应的Base R实现效果:
正如Camille提到的,新版ggplot2支持次轴,但必须基于主轴的转换逻辑——这里的核心是把主轴的累计值转换为次轴的百分比,关键在于手动指定次轴的breaks和labels,让它们和累计值一一对应。下面是修改后的ggplot2代码,完美实现需求:
library(tidyverse) d <- tribble( ~ category, ~defect, "price", 80, "schedule", 27, "supplier", 66, "contact", 94, "item", 33 ) %>% arrange(desc(defect)) %>% mutate( cumsum = cumsum(defect), freq = round(defect / sum(defect), 3), cum_freq = cumsum(freq) ) %>% # 调整分类顺序,让最大的缺陷类别在最左边 mutate(category = fct_reorder(category, defect) %>% fct_rev()) # 准备次轴的断点和标签: # 次轴断点需要是主轴累计值除以总缺陷数(转换为0-1的比例,对应次轴的百分比) sec_breaks <- d$cumsum / max(d$cumsum) # 标签直接用预先计算好的累计频率转成百分比格式 sec_labels <- scales::percent(d$cum_freq) ggplot(d, aes(x = category)) + geom_col(aes(y = defect), fill = "steelblue", alpha = 0.8) + geom_point(aes(y = cumsum), color = "darkorange", size = 2) + geom_line(aes(y = cumsum, group = 1), color = "darkorange", linewidth = 1) + scale_y_continuous( name = "缺陷数量", # 主轴也可以设置断点对应累计值,让图表更清晰 breaks = c(0, d$cumsum), sec.axis = sec_axis( ~ . / max(d$cumsum), name = "累计百分比", breaks = sec_breaks, labels = sec_labels ) ) + labs(title = "优化后的ggplot2帕累托图") + theme_minimal() + theme(axis.text.x = element_text(angle = 45, hjust = 1))
关键逻辑说明:
- 次轴的本质是对主轴数值做线性转换(这里是除以总缺陷数得到比例),所以次轴的
breaks必须是转换后的数值(累计值/总缺陷数) labels直接使用提前计算好的cum_freq,转成百分比格式,这样每个刻度就精准对应到了累计值的位置- 同时主轴也可以设置对应累计值的断点,让整个图表的刻度更直观
内容的提问来源于stack exchange,提问作者stackinator
相关产品推荐
相关产品推荐

