如何用循环遍历星期几简化R语言数据统计与可视化代码?
R语言代码优化方案
需求说明
刚学习R语言两周,现有绘图代码通过硬编码编写了多条annotate语句添加注释,希望优化实现:
- 自动计算
minutesworn列按weekday分组后的均值最大值 - 将该最大值加上指定偏移量后,作为所有注释的统一y轴坐标,替代重复的
annotate语句
现有代码与数据集
数据集
data_df <- tibble::tribble( ~id, ~activitydate, ~totalsteps, ~totaldistance, ~sedentaryminutes, ~calories, ~activeminutes, ~totalminutesasleep, ~totaltimeinbed, ~timeawakeinbed, ~month, ~weekday, ~minutesworn, ~alldaywear, 1503960366, as.Date("2016-04-12"), 13162, 8.5, 728, 1985, 366, 327, 346, 19, "April", "Tuesday", 1094, TRUE, 1503960366, as.Date("2016-04-13"), 10735, 6.96999979019165, 776, 1797, 257, 384, 407, 23, "April", "Wednesday", 1033, TRUE, 1503960366, as.Date("2016-04-15"), 9762, 6.28000020980835, 726, 1745, 272, 412, 442, 30, "April", "Friday", 998, TRUE, 1503960366, as.Date("2016-04-16"), 12669, 8.15999984741211, 773, 1863, 267, 340, 367, 27, "April", "Saturday", 1040, TRUE, 1503960366, as.Date("2016-04-17"), 9705, 6.48000001907349, 539, 1728, 222, 700, 712, 12, "April", "Sunday", 761, TRUE, 1503960366, as.Date("2016-04-19"), 15506, 9.88000011444092, 775, 2035, 345, 304, 320, 16, "April", "Tuesday", 1120, TRUE, 1503960366, as.Date("2016-04-20"), 10544, 6.67999982833862, 818, 1786, 245, 360, 377, 17, "April", "Wednesday", 1063, TRUE, 1503960366, as.Date("2016-04-21"), 9819, 6.34000015258789, 838, 1775, 238, 325, 364, 39, "April", "Thursday", 1076, TRUE, )
现有分析与绘图代码
# 查找均值的最大值 data_df %>% group_by(weekday) %>% summarize(max(mean(minutesworn))) # 绘图 data_df %>% group_by(weekday) %>% summarize(mean_wear = mean(minutesworn)) %>% ggplot(mapping = aes(x = factor(weekday, level = c('Sunday', 'Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday')), y = mean_wear, fill = weekday)) + geom_col() + labs(title = "Minutes Worn by Weekday", caption = "Data Collected in 2016") + xlab("Weekday") + ylab("Average Minutes Worn") + annotate("text", x = "Friday", y = 1052, label = "Friday") + annotate("text", x = "Saturday", y = 1022, label = "Saturday") + annotate("text", x = "Sunday", y = 977, label = "Sunday") + annotate("text", x = "Monday", y = 1040, label = "Monday") + annotate("text", x = "Tuesday", y = 1057, label = "Tuesday") + annotate("text", x = "Wednesday", y = 1010, label = "Wednesday") + annotate("text", x = "Thursday", y = 1008, label = "Thursday")
优化后的代码
# 加载依赖包 library(dplyr) library(ggplot2) # 1. 预处理数据:计算分组均值,设置星期顺序 summary_df <- data_df %>% group_by(weekday) %>% summarize(mean_wear = mean(minutesworn)) %>% mutate(weekday = factor(weekday, level = c('Sunday', 'Monday', 'Tuesday', 'Wednesday', 'Thursday', 'Friday', 'Saturday'))) %>% arrange(weekday) # 2. 计算注释的y轴坐标:均值最大值 + 自定义偏移量(可按需调整) offset <- 30 annotate_y <- max(summary_df$mean_wear) + offset # 3. 绘图并添加注释 ggplot(summary_df, aes(x = weekday, y = mean_wear, fill = weekday)) + geom_col() + # 用geom_text替代重复的annotate,直接映射数据生成注释 geom_text(aes(label = weekday), y = annotate_y) + labs(title = "Minutes Worn by Weekday", caption = "Data Collected in 2016", x = "Weekday", y = "Average Minutes Worn") + # 可选:扩展y轴范围,避免注释超出画布 expand_limits(y = annotate_y + offset/2)
优化亮点
- 消除重复代码:用一行
geom_text替代7条重复的annotate语句,代码更简洁易维护 - 自动计算坐标:无需硬编码y值,通过
max()获取均值最大值,结合偏移量自动生成注释位置 - 保留自定义顺序:提前对
weekday设置因子水平,确保绘图顺序符合需求 - 灵活性高:偏移量
offset可根据可视化需求调整,控制注释与柱子顶部的距离
内容的提问来源于stack exchange,提问作者Cyris Zeiders
相关产品推荐
相关产品推荐

