R语言计算学生考试序列次数报错及结果正确性排查
处理学生考试3位0/1序列统计的解决方案
1. 明确数据结构
假设你的宽表结构如下(示例):
set.seed(123) student_exams <- data.frame( student_id = paste0("S", 1:5), exam1 = sample(0:1, 5, replace = TRUE), exam2 = sample(0:1, 5, replace = TRUE), exam3 = sample(0:1, 5, replace = TRUE), exam4 = sample(0:1, 5, replace = TRUE), exam5 = sample(0:1, 5, replace = TRUE) )
其中student_id是学生标识,后续列对应各次考试的0/1结果。
2. 修复mapply报错问题
你遇到的Error in input[i, ] : incorrect number of dimensions,核心原因是:
- 传入自定义函数的是一维向量,但函数里误用了二维索引(
[i,]); - 或mapply调用逻辑错误,按列而非按行处理学生数据。
正确的处理逻辑是按行提取每个学生的考试向量,再批量处理。
3. 编写序列统计函数
以下函数接收单个学生的考试向量,生成所有连续3位序列,并统计8种可能0/1组合的出现次数(未出现的组合补0):
count_3mers <- function(score_vec) { # 过滤无效值(仅保留0/1) score_vec <- score_vec[score_vec %in% 0:1] n_exams <- length(score_vec) # 考试次数不足3次时返回全0结果 if(n_exams < 3) { return(setNames(rep(0, 8), c("000", "001", "010", "011", "100", "101", "110", "111"))) } # 生成连续3位的字符串序列 mers <- sapply(1:(n_exams - 2), function(i) { paste(score_vec[i:(i+2)], collapse = "") }) # 强制统计所有8种可能序列,确保未出现的计数为0 all_possible <- c("000", "001", "010", "011", "100", "101", "110", "111") mer_counts <- table(factor(mers, levels = all_possible)) return(as.vector(mer_counts)) }
4. 按学生批量处理
用apply按行处理考试数据,得到每个学生的序列计数:
# 提取纯考试成绩列(排除student_id) exam_scores <- student_exams[, -which(colnames(student_exams) == "student_id")] # 按行应用统计函数 student_mer_counts <- apply(exam_scores, 1, count_3mers) # 转换为结构化结果表 result_df <- data.frame( student_id = student_exams$student_id, t(student_mer_counts) )
5. 解决for循环计数过高问题
如果你的for循环结果偏高,大概率是序列生成的索引范围错误:
- n次考试对应的有效连续3位序列数量应为
n-2个(例如5次考试对应3个序列:1-3、2-4、3-5); - 检查循环逻辑,避免错误生成超出索引范围的序列,或误将不同学生的序列合并统计。
6. 计算条件概率
以P(第三位=1 | 前两位=01)为例,对每个学生的计数计算条件概率:
# 添加条件概率列,分母为0时设为NA(无对应序列) result_df$P_1_given_01 <- with(result_df, ifelse((`010` + `011`) == 0, NA, `011` / (`010` + `011`)) )
内容的提问来源于stack exchange,提问作者stats_noob
相关产品推荐
相关产品推荐

