基于循环与条件筛选重构学生测试数据生成答案键
问题:批量生成测试题答案键数据框
我需要重构学生测试数据df_test以生成答案键列表。其中Q1.C至Q20.C表示学生对应题目作答是否正确,Q1.CH至Q20.CH为学生的具体答案。目标输出为包含**Question(题目)和Answer_Key(答案键)**的数据框。
之前手动创建多个子表再合并,20道题时可行,但题目超过100道时手动工作量极大。尝试用循环实现但未成功,现有手动代码如下:
x1<-data.frame(table(df_test$Q1.CH,df_test$Q1.C))%>% subset(Freq!=0 & Var2==1)%>% mutate(Question_ID="Q1")%>% select(Var1,Question_ID) x2<-data.frame(table(df_test$Q2.CH,df_test$Q2.C))%>% subset(Freq!=0 & Var2==1)%>% mutate(Question_ID="Q2")%>% select(Var1,Question_ID) # ... 中间省略x3到x19的代码 ... x20<-data.frame(table(df_test$Q20.CH,df_test$Q20.C))%>% subset(Freq!=0 & Var2==1)%>% mutate(Question_ID="Q20")%>% select(Var1,Question_ID) x<-do.call(rbind.data.frame, mget(ls(pattern = "^x")))
解决方案1:使用tidyverse批量处理(推荐)
这种方法无需手动逐个定义子表,适配任意数量的题目:
library(tidyverse) # 提取所有唯一的题目编号(比如Q1、Q2...Qn) question_ids <- str_extract(names(df_test), "^Q\\d+") %>% unique() %>% na.omit() # 批量处理每个题目,自动合并结果 answer_key_df <- map_dfr(question_ids, function(q_id) { # 拼接对应题目的答案列和对错列名 ch_col <- paste0(q_id, ".CH") c_col <- paste0(q_id, ".C") # 筛选作答正确的记录,提取答案并去重 df_test %>% filter(.data[[c_col]] == 1) %>% select(Answer_Key = .data[[ch_col]]) %>% distinct() %>% mutate(Question = q_id) %>% select(Question, Answer_Key) })
代码说明:
str_extract从列名中自动提取所有Q开头的题目编号map_dfr遍历每个题目,处理后自动合并为一个数据框.data[[col]]用于动态引用列名,避免硬编码distinct()确保每个题目的答案键唯一(重复正确答案只保留一次)
解决方案2:基础R循环实现
如果偏好使用基础R,以下循环代码可以实现同样功能:
# 初始化空结果数据框 answer_key_df <- data.frame( Question = character(), Answer_Key = character(), stringsAsFactors = FALSE ) # 提取所有唯一的题目编号 question_ids <- gsub("\\.(C|CH)$", "", names(df_test)[grepl("^Q\\d+\\.(C|CH)$", names(df_test))]) %>% unique() # 循环处理每个题目 for(q_id in question_ids) { # 获取对应列名 ch_col <- paste0(q_id, ".CH") c_col <- paste0(q_id, ".C") # 筛选作答正确且非NA的答案,去重 correct_answers <- unique(df_test[[ch_col]][df_test[[c_col]] == 1 & !is.na(df_test[[c_col]])]) # 如果存在正确答案,添加到结果数据框 if(length(correct_answers) > 0) { answer_key_df <- rbind(answer_key_df, data.frame(Question = q_id, Answer_Key = correct_answers, stringsAsFactors = FALSE)) } }
代码说明:
grepl筛选出所有题目相关的列,gsub去掉后缀提取题目编号- 循环中动态引用列,筛选正确答案后去重
- 判断
length(correct_answers)避免添加空行
内容的提问来源于stack exchange,提问作者Ksalad
相关产品推荐
相关产品推荐

