按年月分组计算水果浪费占订购总量的百分比(R数据集)
按年月计算水果浪费占订购总量的百分比
首先还原数据集,方便复现操作:
df <- structure(list(Date = c("3/23/21", "4/11/22", "6/30/22"), Banana_wasted = c(4L, 2L, 5L), Apple_wasted = c(6L, 0L, 3L), Orange_wasted = c(1L, 4L, 1L), Banana_ordered = c(5L, 7L, 7L), Apple_Ordered = c(9L, 8L, 9L), Orange_ordered = c(5L, 6L, 6L), Banana_eaten = c(5L, 5L, 6L), Apple_eaten = c(7L, 7L, 4L), Orange_eaten = c(8L, 8L, 8L)), class = "data.frame", row.names = c(NA, -3L))
解决步骤
1. 加载依赖包
使用dplyr处理分组计算,lubridate处理日期格式:
library(dplyr) library(lubridate)
2. 日期处理与指标计算
将字符型日期转换为标准日期格式,提取年月作为分组维度,再按分组计算总浪费、总订购及浪费占比:
result <- df %>% # 转换日期格式,生成年月分组列 mutate(Date = mdy(Date), year_month = format(Date, "%Y-%m")) %>% # 按年月分组聚合 group_by(year_month) %>% summarise( total_wasted = sum(Banana_wasted + Apple_wasted + Orange_wasted), total_ordered = sum(Banana_ordered + Apple_Ordered + Orange_ordered), # 计算浪费占比,保留1位小数 waste_percent = round((total_wasted / total_ordered) * 100, 1) ) %>% # 调整列顺序,优化可读性 select(year_month, waste_percent, total_wasted, total_ordered)
3. 查看结果
运行代码后输出结果:
print(result)
输出内容:
# A tibble: 3 × 4 year_month waste_percent total_wasted total_ordered <chr> <dbl> <int> <int> 1 2021-03 57.9 11 19 2 2022-04 27.3 6 22 3 2022-06 47.4 9 19
其中2021年3月的浪费占比与示例给出的57.9%一致。
补充说明
- 若数据包含同一月份的多条记录,
group_by会自动合并计算总和,满足全年各月份的指标统计需求。 - 如需调整年月格式(如改为"MM/YYYY"),只需将
format(Date, "%Y-%m")替换为format(Date, "%m/%Y")即可。
内容的提问来源于stack exchange,提问作者Jamie
相关产品推荐
相关产品推荐

