ggplot堆叠100%面积图分组填充不显示问题求助
解决堆叠100%面积图分组填充问题
问题根源
你的代码存在两个核心问题:
- 百分比计算逻辑错误:当前
mutate(percentage = n / sum(n))是在group_by(start_hour, member_casual)的分组内执行的,每个分组的sum(n)就是自身的n,导致所有percentage值都是1,图表自然无法显示分层填充效果。 - 未明确指定堆叠模式:即使百分比计算正确,默认的
geom_area也可能因数据结构问题无法正确堆叠,需要显式设置位置参数。
修正方案
方案一:先正确计算百分比再绘图
调整数据处理逻辑,确保每个小时内两类骑行者的百分比总和为1:
start_hour_dist <- clean_trips %>% group_by(start_hour, member_casual) %>% summarise(n = n(), .groups = "drop_last") %>% # 保留start_hour分组用于后续计算 mutate(percentage = n / sum(n)) %>% ungroup()
绘图时指定堆叠模式:
ggplot(start_hour_dist, aes(x = start_hour, y = percentage, fill = member_casual)) + geom_area(position = "stack")
方案二:让ggplot自动计算100%堆叠(更简洁)
无需提前计算百分比,直接用原始计数,通过position = "fill"让ggplot自动处理占比:
# 数据处理仅统计各分组计数 start_hour_dist <- clean_trips %>% group_by(start_hour, member_casual) %>% summarise(n = n(), .groups = "drop") # 绘制100%堆叠面积图并格式化y轴为百分比 ggplot(start_hour_dist, aes(x = start_hour, y = n, fill = member_casual)) + geom_area(position = "fill") + scale_y_continuous(labels = scales::percent)
额外优化建议
- 确保
start_hour是连续数值型变量,如果是因子类型需先转换:start_hour <- as.numeric(as.character(start_hour)) - 若存在某小时缺失某类骑行者的情况,可补全缺失组避免图表断裂:
start_hour_dist <- clean_trips %>% group_by(start_hour, member_casual) %>% summarise(n = n(), .groups = "drop") %>% complete(start_hour = 0:24, member_casual, fill = list(n = 0))
内容的提问来源于stack exchange,提问作者Ben He
相关产品推荐
相关产品推荐

