You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现不依赖member_casual取值的分组百分比直方图?

问题

我有一个member_casual列,最多包含3个取值,希望为每个取值绘制百分比直方图以作对比。关键要求是:百分比计算需基于member_casual对应取值的行数,而非总行数。

我已通过以下代码实现该需求,但需要为每个取值手动添加代码,当member_casual取值变化时必须重写代码,希望找到不依赖具体取值的实现方法:

dataCustomer <- tripDataFiles %>% 
  filter( member_casual == "Customer")
dataSubscriber <- tripDataFiles %>% 
  filter( member_casual == "Subscriber")
dataDependent <- tripDataFiles %>% 
  filter( member_casual == "Dependent")
ggplot(dataCustomer, aes(x=tripduration, y =  stat(count / sum(count))))+
  geom_histogram(aes(fill='customer'), alpha = 0.5)+
  geom_histogram(data=dataSubscriber, aes(fill='subscriber'), alpha = 0.5)+
  geom_histogram(data=dataDependent, aes(fill='dependent'), alpha = 0.5)+
  scale_y_continuous(labels = scales::percent)

补充说明

数据为2015-2017年的骑行数据,已转换为2023年格式:

tripDataFiles17 <- dataFileNames %>%
  grep(x = dataFileNames, pattern = '2017', value = TRUE) %>% #筛选2017年数据
  grep(pattern = 'station',  x = ., ignore.case = TRUE, invert = TRUE, value = TRUE) %>% #移除站点相关文件
  lapply(fread) %>% #读取选中文件
  rbindlist() %>%  #合并数据
  rename(
    started_at = start_time,
    ended_at = end_time
  ) %>% 
  mutate(started_at = parse_date_time(started_at,dateTimeFormat), ended_at = parse_date_time(ended_at,dateTimeFormat)) #转换时间格式

tripDataFiles <- rbindlist( list(tripDataFiles15_16, tripDataFiles17)) %>%
  rename(
    ride_id = trip_id,
    start_station_id = from_station_id,
    start_station_name = from_station_name,
    end_station_id = to_station_id,
    end_station_name = to_station_name,
    member_casual = usertype
  )

数据示例:

dput(tripDataFiles[1:20, c("member_casual", "tripduration")])
structure(list(member_casual = c("Subscriber", "Customer", "Subscriber", 
"Customer", "Subscriber", "Subscriber", "Subscriber", "Subscriber", 
"Subscriber", "Customer", "Customer", "Customer", "Customer", 
"Subscriber", "Subscriber", "Subscriber", "Subscriber", "Subscriber", 
"Subscriber", "Subscriber"), tripduration = c(299L, 940L, 751L, 
1240L, 1292L, 175L, 930L, 383L, 260L, 1123L, 1167L, 231L, 1092L, 
585L, 401L, 177L, 653L, 303L, 223L, 353L)), row.names = c(NA, 
-20L), class = c("data.table", "data.frame"), ...)

解决方案

可以通过两种方式实现不依赖member_casual具体取值的百分比直方图:

方法1:直接用ggplot分组计算

利用ggplot的分组功能,结合after_stat()自动按组计算百分比,无需拆分数据:

ggplot(tripDataFiles, aes(x = tripduration, y = after_stat(count / sum(count)), fill = member_casual)) +
  geom_histogram(position = "identity", alpha = 0.5) +
  scale_y_continuous(labels = scales::percent) +
  labs(fill = "用户类型") # 可选:自定义图例标题
  • 核心逻辑:fill = member_casual自动按用户类型分组,after_stat(count / sum(count))会在每个分组内独立计算百分比,完全匹配需求;
  • position = "identity"保证直方图重叠显示,和原代码效果一致。

方法2:先预处理数据(更灵活)

如果需要自定义分组逻辑或对百分比做额外处理,可先通过dplyr预处理数据再绘图:

library(dplyr)

# 预处理:按用户类型分组计算每组内的百分比
tripData_processed <- tripDataFiles %>%
  group_by(member_casual) %>%
  mutate(group_total = n()) %>% # 记录每组的总行数
  group_by(member_casual, tripduration) %>%
  summarise(count = n(), group_total = first(group_total), .groups = "drop") %>%
  mutate(percent = count / group_total)

# 绘制直方图
ggplot(tripData_processed, aes(x = tripduration, y = percent, fill = member_casual)) +
  geom_histogram(stat = "identity", position = "identity", alpha = 0.5, bins = 30) + # 可自定义bins数量
  scale_y_continuous(labels = scales::percent) +
  labs(fill = "用户类型")
  • 这种方式能更直观地控制数据处理流程,方便后续调整直方图区间或添加其他统计逻辑。

内容的提问来源于stack exchange,提问作者Arthur. R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 02:07:21