如何在R中基于学生、测试名称与日期创建测试尝试次数索引变量?
问题:基于学生ID和测试名称生成尝试次数索引变量
我在R语言中尝试基于学生ID、测试名称和测试日期创建一个索引变量。数据包含学生重复参加同一测试的记录,每个观测对应不同分数,需要按日期排序,给每个学生的每类测试标记尝试次数——同一学生同一测试的计数从1开始递增。
原始数据
student <- c(1,1,1,1,1,1,2,2,2,3,3,3,3,3) test <- c("math","math","reading","math","reading","reading","reading","math","reading","math","math","math","reading","reading") date <- c(1,2,3,3,4,5,2,3,5,1,2,3,4,5) data <- data.frame(student,test,date) print(data)
输出结果:
student test date 1 1 math 1 2 1 math 2 3 1 reading 3 4 1 math 3 5 1 reading 4 6 1 reading 5 7 2 reading 2 8 2 math 3 9 2 reading 5 10 3 math 1 11 3 math 2 12 3 math 3 13 3 reading 4 14 3 reading 5
期望结果
希望添加表示学生对应测试尝试次数的id变量,最终数据如下:
student test date id 1 1 math 1 1 2 1 math 2 2 3 1 reading 3 1 4 1 math 3 3 5 1 reading 4 2 6 1 reading 5 3 7 2 reading 2 1 8 2 math 3 1 9 2 reading 5 2 10 3 math 1 1 11 3 math 2 2 12 3 math 3 3 13 3 reading 4 1 14 3 reading 5 2
我尝试过的代码
我之前尝试用cumsum实现,但无法在新分组时重置计数:
tests <- transform(tests, ID = as.numeric(factor(EMPLID))) tests$id <- cumsum(!duplicated(tests[1:3]))
解决方案
方法1:基础R(无需额外包)
使用ave()函数实现分组计数,结合rank()处理日期排序,ties.method = "first"确保同日期的记录按原始顺序递增:
# 生成id变量 data$id <- ave(data$date, data$student, data$test, FUN = function(x) rank(x, ties.method = "first")) # 查看结果 print(data)
方法2:tidyverse包(更直观)
使用dplyr的分组和排序功能,逻辑更清晰:
library(dplyr) data <- data %>% # 按学生和测试分组 group_by(student, test) %>% # 组内按日期排序 arrange(date, .by_group = TRUE) %>% # 生成组内递增的id mutate(id = row_number()) %>% # 取消分组 ungroup() # 查看结果 print(data)
两种方法都能得到你期望的结果,基础R方法无需安装额外包,tidyverse方法代码可读性更强,适合处理复杂数据操作。
内容的提问来源于stack exchange,提问作者candace
相关产品推荐
相关产品推荐

