如何在R语言中用循环计算每人的物品浪费占比?
解决方案
首先还原你的数据集:
df <- structure(list(Person= c("John Smith", "John Smith", "John Smith", "John Smith", "John Smith", "John Smith", "John Smith", "Martin Harris", "Martin Harris", "Martin Harris", "Kyle Short"), Item.Order = c("ABC", "ABC", "DEF", "ABC", "IJK", "ABC", "DEF", "IJK", "ABC", "ABC", "DEF"), Status = c("R", "W", "R", "R", "W", "W", "W", "R", "W", "R", "W")), class = "data.frame", row.names = c(NA, -11L))
方法一:用循环实现(符合你的需求)
# 获取所有唯一的人员 unique_people <- unique(df$Person) # 初始化结果数据框 waste_ratio <- data.frame(Person = character(), Ratio = numeric(), stringsAsFactors = FALSE) # 循环每个人员计算比例 for(person in unique_people) { # 筛选当前人员的所有记录 person_data <- df[df$Person == person, ] # 总操作次数 total_ops <- nrow(person_data) # 浪费的次数(Status为"W"的记录数) waste_count <- sum(person_data$Status == "W") # 计算比例 ratio <- waste_count / total_ops # 加入结果框 waste_ratio <- rbind(waste_ratio, data.frame(Person = person, Ratio = ratio)) } # 查看结果 print(waste_ratio)
运行后会得到每个人的浪费比例,比如John Smith的结果就是0.5714286。
方法二:更高效的非循环实现(R语言推荐写法)
如果不需要强制用循环,用以下方法会更简洁高效,处理大数据集时速度更快:
用dplyr包:
library(dplyr) waste_ratio <- df %>% group_by(Person) %>% summarise( Ratio = sum(Status == "W") / n() ) %>% ungroup() print(waste_ratio)
用base R的aggregate函数:
waste_ratio <- aggregate(Status ~ Person, data = df, function(x) sum(x == "W") / length(x)) colnames(waste_ratio)[2] <- "Ratio" print(waste_ratio)
这两种非循环方法结果和循环完全一致,但代码更简洁。
内容的提问来源于stack exchange,提问作者user16368421
相关产品推荐
相关产品推荐

