如何实现不依赖member_casual取值的分组百分比直方图?
问题
我有一个member_casual列,最多包含3个取值,希望为每个取值绘制百分比直方图以作对比。关键要求是:百分比计算需基于member_casual对应取值的行数,而非总行数。
我已通过以下代码实现该需求,但需要为每个取值手动添加代码,当member_casual取值变化时必须重写代码,希望找到不依赖具体取值的实现方法:
dataCustomer <- tripDataFiles %>% filter( member_casual == "Customer") dataSubscriber <- tripDataFiles %>% filter( member_casual == "Subscriber") dataDependent <- tripDataFiles %>% filter( member_casual == "Dependent") ggplot(dataCustomer, aes(x=tripduration, y = stat(count / sum(count))))+ geom_histogram(aes(fill='customer'), alpha = 0.5)+ geom_histogram(data=dataSubscriber, aes(fill='subscriber'), alpha = 0.5)+ geom_histogram(data=dataDependent, aes(fill='dependent'), alpha = 0.5)+ scale_y_continuous(labels = scales::percent)
补充说明
数据为2015-2017年的骑行数据,已转换为2023年格式:
tripDataFiles17 <- dataFileNames %>% grep(x = dataFileNames, pattern = '2017', value = TRUE) %>% #筛选2017年数据 grep(pattern = 'station', x = ., ignore.case = TRUE, invert = TRUE, value = TRUE) %>% #移除站点相关文件 lapply(fread) %>% #读取选中文件 rbindlist() %>% #合并数据 rename( started_at = start_time, ended_at = end_time ) %>% mutate(started_at = parse_date_time(started_at,dateTimeFormat), ended_at = parse_date_time(ended_at,dateTimeFormat)) #转换时间格式 tripDataFiles <- rbindlist( list(tripDataFiles15_16, tripDataFiles17)) %>% rename( ride_id = trip_id, start_station_id = from_station_id, start_station_name = from_station_name, end_station_id = to_station_id, end_station_name = to_station_name, member_casual = usertype )
数据示例:
dput(tripDataFiles[1:20, c("member_casual", "tripduration")]) structure(list(member_casual = c("Subscriber", "Customer", "Subscriber", "Customer", "Subscriber", "Subscriber", "Subscriber", "Subscriber", "Subscriber", "Customer", "Customer", "Customer", "Customer", "Subscriber", "Subscriber", "Subscriber", "Subscriber", "Subscriber", "Subscriber", "Subscriber"), tripduration = c(299L, 940L, 751L, 1240L, 1292L, 175L, 930L, 383L, 260L, 1123L, 1167L, 231L, 1092L, 585L, 401L, 177L, 653L, 303L, 223L, 353L)), row.names = c(NA, -20L), class = c("data.table", "data.frame"), ...)
解决方案
可以通过两种方式实现不依赖member_casual具体取值的百分比直方图:
方法1:直接用ggplot分组计算
利用ggplot的分组功能,结合after_stat()自动按组计算百分比,无需拆分数据:
ggplot(tripDataFiles, aes(x = tripduration, y = after_stat(count / sum(count)), fill = member_casual)) + geom_histogram(position = "identity", alpha = 0.5) + scale_y_continuous(labels = scales::percent) + labs(fill = "用户类型") # 可选:自定义图例标题
- 核心逻辑:
fill = member_casual自动按用户类型分组,after_stat(count / sum(count))会在每个分组内独立计算百分比,完全匹配需求; position = "identity"保证直方图重叠显示,和原代码效果一致。
方法2:先预处理数据(更灵活)
如果需要自定义分组逻辑或对百分比做额外处理,可先通过dplyr预处理数据再绘图:
library(dplyr) # 预处理:按用户类型分组计算每组内的百分比 tripData_processed <- tripDataFiles %>% group_by(member_casual) %>% mutate(group_total = n()) %>% # 记录每组的总行数 group_by(member_casual, tripduration) %>% summarise(count = n(), group_total = first(group_total), .groups = "drop") %>% mutate(percent = count / group_total) # 绘制直方图 ggplot(tripData_processed, aes(x = tripduration, y = percent, fill = member_casual)) + geom_histogram(stat = "identity", position = "identity", alpha = 0.5, bins = 30) + # 可自定义bins数量 scale_y_continuous(labels = scales::percent) + labs(fill = "用户类型")
- 这种方式能更直观地控制数据处理流程,方便后续调整直方图区间或添加其他统计逻辑。
内容的提问来源于stack exchange,提问作者Arthur. R
相关产品推荐
相关产品推荐

