按Subjectid统计响应值出现次数并生成带前缀计数列(R语言)
按subjectid分组统计响应次数并拼接变量
初始数据
subjectid <- c(1, 1, 1, 2, 2, 3, 3, 3, 4, 4, 5) response <- c("PD", "PD", "SD", "PD", "SD", "PD", "SD", "SD", "SD", "PD", "PR") df <- data.frame(subjectid, response)
实现方法
方法1:使用dplyr包
通过分组生成组内序号,再拼接变量:
library(dplyr) df_result <- df %>% # 按subjectid和response分组 group_by(subjectid, response) %>% # 生成当前response在对应subject内的出现次数 mutate(count = row_number(), # 拼接响应值与计数 response_count = paste0(response, count)) %>% # 取消分组 ungroup() print(df_result)
方法2:基础R实现(无需额外包)
利用ave函数完成分组计数:
# 生成每个subject内各response的出现次数 df$count <- ave(rep(1, nrow(df)), df$subjectid, df$response, FUN = seq_along) # 拼接成新变量 df$response_count <- paste0(df$response, df$count) print(df)
输出结果
subjectid response count response_count 1 1 PD 1 PD1 2 1 PD 2 PD2 3 1 SD 1 SD1 4 2 PD 1 PD1 5 2 SD 1 SD1 6 3 PD 1 PD1 7 3 SD 1 SD1 8 3 SD 2 SD2 9 4 SD 1 SD1 10 4 PD 1 PD1 11 5 PR 1 PR1
内容的提问来源于stack exchange,提问作者samthecodegod
相关产品推荐
相关产品推荐

