如何用ggplot2在R中可视化含团队、地点及失败原因的数据集?
解决方案:ggplot2 可视化团队-地点-失败原因关联
1. 数据格式转换(关键前提)
你的数据集里reason_type1到reason_type4是宽格式结构,ggplot2对长格式数据的支持更友好,先将这部分转换为长格式:
library(tidyverse) # 假设你的数据集名为 df df_long <- df %>% pivot_longer( cols = starts_with("reason_type"), names_to = "failure_reason", values_to = "reason_ratio" ) %>% # 计算各失败原因对应的实际天数 mutate(reason_fail_days = fail_days * reason_ratio)
2. 核心可视化代码(两种适配方案)
方案一:堆叠柱状图(团队分组+地点配色+失败原因堆叠)
适合直观对比同一团队在不同地点下,各失败原因的天数分布:
ggplot(df_long, aes(x = Group, y = reason_fail_days, fill = failure_reason, color = locat)) + geom_col(position = "stack", alpha = 0.8, size = 1) + # 可选:添加失败天数标签,提升可读性 geom_text(aes(label = round(reason_fail_days, 1)), position = position_stack(vjust = 0.5), size = 3) + labs( title = "各团队不同地点的失败天数及原因分布", x = "参赛团队", y = "失败天数", fill = "失败原因", color = "比赛地点" ) + theme_minimal() + theme( plot.title = element_text(hjust = 0.5, size = 14, face = "bold"), axis.text.x = element_text(angle = 45, hjust = 1) )
方案二:分面柱状图(按地点拆分展示)
如果需要更清晰地按地点维度拆分查看各团队数据,分面布局会更合适:
ggplot(df_long, aes(x = Group, y = reason_fail_days, fill = failure_reason)) + geom_col(position = "stack", alpha = 0.8) + geom_text(aes(label = round(reason_fail_days, 1)), position = position_stack(vjust = 0.5), size = 3) + facet_wrap(~locat, ncol = 2) # 可根据地点数量调整列数 labs( title = "各地点下团队的失败天数及原因分布", x = "参赛团队", y = "失败天数", fill = "失败原因" ) + theme_minimal() + theme( plot.title = element_text(hjust = 0.5, size = 14, face = "bold"), axis.text.x = element_text(angle = 45, hjust = 1), strip.text = element_text(face = "bold", size = 10) )
3. 可视化优化建议
- 配色优化:用专业配色替换默认色,比如
scale_fill_brewer(palette = "Set2")设置失败原因的填充色,scale_color_viridis_d()设置地点的边框色,避免颜色冲突。 - 数据简化:如果某类失败原因占比极低(比如<5%),可合并为“其他”类别,避免图表拥挤。
- 交互增强:结合
plotly包将静态图转为交互式图表,点击可查看单条数据详情:library(plotly) ggplotly(your_ggplot_object) - 辅助信息:添加各团队总失败天数标注,或绘制平均失败天数参考线,强化对比逻辑。
- 标签调整:团队名称过长时可缩短或换行,调整文字角度和大小,避免重叠。
内容的提问来源于stack exchange,提问作者Ollie
相关产品推荐
相关产品推荐

