R ggplot2多线图绘制:无需单独数据框显示百分比
问题
想用R的ggplot2绘制多线图,追踪两个地点(L1、L2)在6个时间点(t1-t6)的人员占比情况。目前能生成双线条图形,但Y轴无法正确显示百分比;现有可行方案需要先创建包含占比的单独数据框,请问有没有更直接的实现方法?
原始数据示例
time <- c("t1", "t1", "t1", "t1", "t1", "t1", "t2", "t2", "t2", "t2", "t2", "t2", "t3", "t3", "t3", "t3", "t3", "t3", "t4", "t4", "t4", "t4", "t4", "t4", "t5", "t5", "t5", "t5", "t5", "t5", "t6", "t6", "t6", "t6", "t6", "t6") location <- c ("L1", "L1", "L1", "L1", "L2", "L2", "L1", "L2", "L2", "L2", "L2", "L2", "L1", "L1", "L1", "L2", "L2", "L2", "L2", "L2", "L2", "L2", "L2", "L2", "L1", "L1", "L2", "L2", "L2", "L2", "L1", "L1", "L2", "L2", "L1", "L2") data <- data.frame (time, location) data table (data$time, exclude = FALSE) table (data$location, exclude = FALSE) table (data$location, data$time, exclude = FALSE)
当前无法显示百分比的代码
ggplot (data = data, mapping = aes ( x = time, y = location, group = location, color = location )) + geom_point (stat = "identity", size = 3) + geom_line (stat = "identity") + ggtitle("各时间点不同地点的人员占比") + xlab("时间") + ylab ("百分比") + coord_cartesian( ylim = c(0, 100))
期望效果的测试代码(需单独数据框)
time <- c("t1", "t1", "t2", "t2", "t3", "t3", "t4", "t4", "t5", "t5", "t6", "t6") location <- c("L1", "L2", "L1", "L2", "L1", "L2", "L1", "L2", "L1", "L2", "L1", "L2") percent <- c(67, 33, 24, 29, 35, 45, 54, 56, 72, 91, 83, 23) test <- data.frame (time, location, percent) ggplot (data = test, mapping = aes ( x = time, y = percent, group = location, color = location )) + # scale_y_continuous(labels = scales::percent) + geom_point (stat = "identity", size = 3) + geom_line (stat = "identity") + ggtitle("月度人员流向家庭/医院的占比图") + xlab("时间") + ylab ("百分比") + coord_cartesian( ylim = c(0, 100))
解决方案
不需要单独创建占比数据框,以下两种方法可直接实现需求:
方法1:利用ggplot的stat_count直接计算占比
通过stat_count的内置变量..prop..计算每个时间点内各地点的占比,再结合scales包格式化Y轴为百分比显示:
library(ggplot2) library(scales) ggplot(data = data, aes(x = time, color = location, group = location)) + stat_count(aes(y = ..prop.. * 100), geom = "line", position = "identity") + stat_count(aes(y = ..prop.. * 100), geom = "point", size = 3, position = "identity") + scale_y_continuous(labels = percent_format(accuracy = 1)) + ggtitle("各时间点不同地点的人员占比") + xlab("时间") + ylab("百分比") + coord_cartesian(ylim = c(0, 100))
..prop..代表每组内的比例值,乘以100转换为百分比数值;position = "identity"确保计算逻辑是每个时间点内各地点的占比,而非全局占比。
方法2:在ggplot调用中直接用dplyr处理数据
如果习惯用dplyr做数据汇总,可直接在ggplot()的data参数里通过管道完成占比计算,无需单独保存中间数据框:
library(ggplot2) library(dplyr) library(scales) ggplot(data = data %>% group_by(time, location) %>% summarise(count = n(), .groups = "drop") %>% group_by(time) %>% mutate(percent = (count / sum(count)) * 100), aes(x = time, y = percent, color = location, group = location)) + geom_point(size = 3) + geom_line() + scale_y_continuous(labels = percent_format(accuracy = 1)) + ggtitle("各时间点不同地点的人员占比") + xlab("时间") + ylab("百分比") + coord_cartesian(ylim = c(0, 100))
这段代码先按「时间+地点」分组统计人数,再按「时间」分组计算每个地点的占比,直接将处理后的结果传入ggplot,一步完成数据处理与绘图。
内容的提问来源于stack exchange,提问作者Eric Boorman
相关产品推荐
相关产品推荐

