R语言帕累托图(Pareto Chart)累计百分比显示异常问题求助
帕累托图累计百分比异常波动问题排查与修复
我在用R绘制软件仓库提交数的帕累托图,验证开发者提交是否符合帕累托法则,但图表里的累计百分比线出现了莫名波动——控制台打印的累计数值明明是正常递增的,图表却乱了。代码和问题现象如下:
我的代码
# 生成帕累托图并保存为PDF的函数 create_pareto_chart <- function(input_file, output_dir) { # 读取CSV文件 data <- read.csv(input_file) # 从文件路径提取项目名称 project_name <- basename(input_file) project_name <- sub("\\.csv$", "", project_name) # 统计每个作者的提交次数 commit_counts <- data %>% count(author) %>% arrange(desc(n)) # 过滤掉空值或缺失的作者 commit_counts <- commit_counts %>% filter(!is.na(author) & author != "") # 计算提交数的累计百分比 total_commits <- sum(commit_counts$n) commit_counts <- commit_counts %>% mutate(cumulative_sum = cumsum(n), cumulative_percentage = 100 * cumulative_sum / total_commits) # 打印累计百分比数值 print(commit_counts) # 创建帕累托图 p <- ggplot(commit_counts, aes(x = reorder(author, -n), y = n)) + geom_bar(stat = "identity", fill = "orange") + geom_line(aes(y = cumulative_percentage * max(n) / 100, group = 1), color = "red") + geom_point(aes(y = cumulative_percentage * max(n) / 100), color = "red") + geom_text(aes(y = cumulative_percentage * max(n) / 100, label = round(cumulative_percentage, 1)), vjust = -0.5, size = 3, color = "blue") + # 为累计线添加标签 geom_hline(yintercept = max(commit_counts$n) * 0.8, color = "red", linetype = "dashed") + scale_y_continuous( sec.axis = sec_axis(~ . * 100 / max(commit_counts$n), name = "提交数累计百分比") ) + labs(title = paste("各作者提交数帕累托图(", project_name, ")", sep = ""), x = "作者(已匿名)", y = "提交次数") + theme(axis.text.x = element_text(angle = 90, hjust = 1)) # 将图表保存为PDF pdf(file = file.path(output_dir, paste(project_name, ".pdf", sep = "")), width = 14, height = 7) print(p) dev.off() }
问题现象
控制台打印的累计百分比是正常递增的:
但图表里的累计线却出现了波动:
问题根源
你用reorder(author, -n)在ggplot的aes里动态排序x轴,但这个排序的稳定性有问题——当多个作者提交数相同时,reorder会打乱你之前用arrange(desc(n))排好的顺序,导致绘图时的x轴元素顺序和commit_counts数据框的行顺序不匹配,累计线的点就会连错,出现波动。
修复方案
1. 提前固定作者的排序顺序
不要在ggplot里动态排序,而是在数据预处理阶段就把author转换成固定顺序的因子,确保x轴顺序和数据行顺序完全一致:
修改数据统计部分的代码:
# 统计每个作者的提交次数、排序、过滤空值,最后固定因子顺序 commit_counts <- data %>% count(author) %>% arrange(desc(n)) %>% filter(!is.na(author) & author != "") %>% mutate(author = factor(author, levels = author)) # 用当前行的author顺序作为因子水平
然后在ggplot的aes里直接用x = author:
p <- ggplot(commit_counts, aes(x = author, y = n)) +
2. 优化累计线的y轴映射(可选)
你现在手动转换累计百分比到左侧y轴的数值范围,容易出错,不如直接用右侧轴的原始百分比数值,让ggplot自动处理映射:
修改累计百分比计算和绘图代码:
# 累计百分比计算保留原始值 commit_counts <- commit_counts %>% mutate(cumulative_sum = cumsum(n), cumulative_percentage = 100 * cumulative_sum / total_commits) # 绘图时调整累计线的y映射 p <- ggplot(commit_counts, aes(x = author, y = n)) + geom_bar(stat = "identity", fill = "orange") + # 直接用累计百分比作为y值,交给右侧轴处理 geom_line(aes(y = cumulative_percentage, group = 1), color = "red") + geom_point(aes(y = cumulative_percentage), color = "red") + geom_text(aes(y = cumulative_percentage, label = round(cumulative_percentage, 1)), vjust = -0.5, size = 3, color = "blue") + scale_y_continuous( name = "提交次数", # 右侧轴直接映射原始百分比 sec.axis = sec_axis(~ ., name = "提交数累计百分比") ) + # 帕累托法则的基准线应该是右侧轴的80%,不是左侧轴的80%最大提交数 geom_hline(yintercept = 80, color = "darkred", linetype = "dashed") + labs(title = paste("各作者提交数帕累托图(", project_name, ")", sep = ""), x = "作者(已匿名)") + theme(axis.text.x = element_text(angle = 90, hjust = 1))
这样不仅能解决波动问题,还能让帕累托基准线更准确(对应80%的累计提交占比)。
内容的提问来源于stack exchange,提问作者Giammaria GIORDANO
相关产品推荐
相关产品推荐

