使用ggplot绘制柱状图时误差线位置错误的原因排查
ggplot柱状图误差线显示异常的解决方法
问题原因
你的数据框是分组的9行数据,每个variable对应3行完全重复的统计值(gfp_mean、ymax、ymin完全一致),导致geom_errorbar在同一个x位置重复绘制3次误差线,视觉上出现重叠偏移;同时分组数据框的默认分组逻辑也会干扰绘图定位。
解决方法
方法1:用去重后的统计数据绘图
直接保留每个variable的唯一统计行,避免重复绘制:
# 转换为普通数据框并去重 gfp_long_uniq <- gfp_long %>% ungroup() %>% distinct(variable, .keep_all = TRUE) # 绘图 ggplot(gfp_long_uniq) + geom_bar(aes(x = variable, y = gfp_mean), stat = "identity", fill = "skyblue", alpha = 0.7) + geom_errorbar(aes(x = variable, ymin = ymin, ymax = ymax), width = 0.4, colour = "orange", alpha = 0.9, size = 1.3)
方法2:重新计算统计量(更规范)
从原始数据直接分组计算均值和误差,生成仅含统计结果的3行数据框,彻底避免数据冗余:
# 计算统计量 gfp_stats <- gfp_long %>% ungroup() %>% group_by(variable) %>% summarise( gfp_mean = mean(gfp), stdev = sd(gfp), ymax = gfp_mean + stdev, ymin = gfp_mean - stdev ) # 绘图 ggplot(gfp_stats) + geom_bar(aes(x = variable, y = gfp_mean), stat = "identity", fill = "skyblue", alpha = 0.7) + geom_errorbar(aes(x = variable, ymin = ymin, ymax = ymax), width = 0.4, colour = "orange", alpha = 0.9, size = 1.3)
方法3:临时取消分组(快速修复)
不想修改数据的话,可在绘图时临时取消分组,并指定只绘制唯一值:
ggplot(ungroup(gfp_long)) + geom_bar(aes(x = variable, y = gfp_mean), stat = "identity", fill = "skyblue", alpha = 0.7) + geom_errorbar(aes(x = variable, ymin = ymin, ymax = ymax), width = 0.4, colour = "orange", alpha = 0.9, size = 1.3, stat = "unique")
原始问题相关信息
异常图表

原始数据
# A tibble: 9 × 6 # Groups: variable [3] variable gfp stdev gfp_mean ymax ymin <chr> <dbl> <dbl> <dbl> <dbl> <dbl> 1 gfp_sums 4286898 466912. 4478746 4945658. 4011834. 2 nc_sums 664845 4378. 662308. 666686. 657930. 3 media_sums 778269 29403. 744335 773738. 714932. 4 gfp_sums 5011021 466912. 4478746 4945658. 4011834. 5 nc_sums 657253 4378. 662308. 666686. 657930. 6 media_sums 726416 29403. 744335 773738. 714932. 7 gfp_sums 4138319 466912. 4478746 4945658. 4011834. 8 nc_sums 664827 4378. 662308. 666686. 657930. 9 media_sums 728320 29403. 744335 773738. 714932.
原始绘图代码
ggplot(gfp_long)+ geom_bar( aes(x=variable, y=gfp_mean), stat="identity", fill="skyblue", alpha=0.7) + geom_errorbar( aes(x=variable, ymin=ymin, ymax=ymax), width=0.4, colour="orange", alpha=0.9, size=1.3)
数据框dput内容
structure(list(variable = c("gfp_sums", "nc_sums", "media_sums", "gfp_sums", "nc_sums", "media_sums", "gfp_sums", "nc_sums", "media_sums" ), gfp = c(4286898, 664845, 778269, 5011021, 657253, 726416, 4138319, 664827, 728320), stdev = c(466911.593911524, 4378.05634195511, 29403.1217900413, 466911.593911524, 4378.05634195511, 29403.1217900413, 466911.593911524, 4378.05634195511, 29403.1217900413), gfp_mean = c(4478746, 662308.333333333, 744335, 4478746, 662308.333333333, 744335, 4478746, 662308.333333333, 744335), ymax = c(4945657.59391152, 666686.389675288, 773738.121790041, 4945657.59391152, 666686.389675288, 773738.121790041, 4945657.59391152, 666686.389675288, 773738.121790041 ), ymin = c(4011834.40608848, 657930.276991378, 714931.878209959, 4011834.40608848, 657930.276991378, 714931.878209959, 4011834.40608848, 657930.276991378, 714931.878209959)), class = c("grouped_df", "tbl_df", "tbl", "data.frame"), row.names = c(NA, -9L), groups = structure(list( variable = c("gfp_sums", "media_sums", "nc_sums"), .rows = structure(list( c(1L, 4L, 7L), c(3L, 6L, 9L), c(2L, 5L, 8L)), ptype = integer(0), class = c("vctrs_list_of", "vctrs_vctr", "list"))), class = c("tbl_df", "tbl", "data.frame" ), row.names = c(NA, -3L), .drop = TRUE))
内容的提问来源于stack exchange,提问作者Salt_Industry
相关产品推荐
相关产品推荐

