R语言遍历DataFrame行实现学生成绩等级自动标注的实现方案
逐行循环(Rowwise Loops)
原始数据集构造
name <- c("John", "Rachel", "Judy", "James", "Oloo") english <- c("70", "50e", "19c", "38^", "33^") math <- c("65", "25c", "68", "32^", "50") science <- c("45", "50", "25c", "27e", "72") social <- c("56", "76", "42", "23^", "68") marks <- data.frame(name, english, math, science, social, stringsAsFactors = FALSE)
等级判定规则
Pass:学生所有成绩无e或^后缀,且全部科目得分≥40Supplementary:学生成绩存在^或e后缀;^代表得分低于40,e代表未参加平时考核,即使得分≥40也判定不合格Special:学生成绩存在c后缀,但无e或^后缀Supplementary & Special:学生成绩同时存在e/^后缀与c后缀Discontinued:学生成绩中e、^后缀累计出现次数≥4次
注意:Discontinued判定优先级最高,其余规则按从上到下优先级判定
可运行实现代码
基础R实现(无需额外依赖)
# 逐行遍历成绩列,生成comment标注 marks$comment <- apply(marks[, -1], 1, function(row_scores) { # 统计e、^后缀出现次数 cnt_e_caret <- sum(grepl("[e^]", row_scores)) # 检查是否存在c后缀 has_c_suffix <- any(grepl("c", row_scores)) # 按优先级判定等级 if (cnt_e_caret >= 4) { return("Discontinued") } if (cnt_e_caret > 0 && has_c_suffix) { return("Supplementary & Special") } if (cnt_e_caret > 0) { return("Supplementary") } if (has_c_suffix) { return("Special") } # 无特殊后缀时提取纯数字分数,检查是否全部≥40 pure_scores <- as.numeric(gsub("[^0-9]", "", row_scores)) if (all(pure_scores >= 40)) { return("Pass") } return(NA_character_) }) # 输出最终结果 print(marks)
dplyr逐行实现(适合tidyverse用户)
library(dplyr) marks <- marks %>% rowwise() %>% mutate( cnt_e_caret = sum(grepl("[e^]", c_across(english:social))), has_c = any(grepl("c", c_across(english:social))), comment = case_when( cnt_e_caret >=4 ~ "Discontinued", cnt_e_caret >0 & has_c ~ "Supplementary & Special", cnt_e_caret >0 ~ "Supplementary", has_c ~ "Special", all(as.numeric(gsub("[^0-9]", "", c_across(english:social))) >=40) ~ "Pass", TRUE ~ NA_character_ ) ) %>% select(-cnt_e_caret, -has_c) %>% ungroup() print(marks)
输出结果验证
运行上述代码后得到的结果与示例完全一致:
name english math science social comment 1 John 70 65 45 56 Pass 2 Rachel 50e 25c 50 76 Supplementary & Special 3 Judy 19c 68 25c 42 Special 4 James 38^ 32^ 27e 23^ Discontinued 5 Oloo 33^ 50 72 68 Supplementary
内容的提问来源于stack exchange,提问作者John Karuitha
相关产品推荐
相关产品推荐

