R语言:提取满足时间≤4年的唯一ID对应最大时间数据集
问题需求
需要生成一个新数据集,满足以下条件:
- 提取原数据中的唯一ID
- 对每个ID,仅保留时间≤4年的记录里时间值最大的那一行,同时保留对应的
status和cancer变量
原始数据结构
data <- structure(list(State = structure(c(1L, 1L, 1L, 1L,1L, 1L, 2L, 2L, 2L, 2L, 3L, 3L, 3L, 3L,3L, 3L, 3L, 3L, 3L), .Label = c("1", "2", "3"), class = "factor"), Time = structure(1:18, .Label = c("0", "1", "2", "3", "4", "5", "0", "1", "2", "3", "0", "1", "2", "3", "4", "5", "6", "7"), class = "factor"), Status = c(0L, 0L, 0L, 0L, 1L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 1L ), cancer = structure(1:18, .Label = c("1", "1", "1", "1", "1", "1", "2", "2", "2", "2", "1", "1", "1", "1", "1", "1", "1", "1"), class = "factor")), .Names = c("ID", "timeYears", "status", "cancer"), class = "data.frame", row.names = c(NA, -18L))
目标数据结构
data1 <- structure(list(State = structure(c(1L, 2L, 3L), .Label = c("1", "2", "3"), class = "factor"), Time = structure(1:3, .Label = c("4", "3", "4"), class = "factor"), Status = c( 1L, 0L, 0L), cancer = structure(1:3, .Label = c("1", "2", "1"), class = "factor")), .Names = c("ID", "timeYears", "status", "cancer"), class = "data.frame", row.names = c(NA, -3L))
解决方案
方法1:使用dplyr包(代码简洁直观)
先将timeYears从因子转换为数值型,再按ID分组筛选目标行:
library(dplyr) result <- data %>% mutate(timeYears = as.numeric(as.character(timeYears))) %>% filter(timeYears <= 4) %>% group_by(ID) %>% filter(timeYears == max(timeYears)) %>% ungroup() %>% select(ID, timeYears, status, cancer)
方法2:使用基础R(无需额外安装包)
通过分组遍历实现需求:
# 转换时间列类型 data$timeYears <- as.numeric(as.character(data$timeYears)) # 筛选时间≤4的记录 filtered_data <- data[data$timeYears <= 4, ] # 分组提取每组最大时间对应的行 result <- do.call(rbind, lapply(split(filtered_data, filtered_data$ID), function(group) { group[group$timeYears == max(group$timeYears), ] })) # 重置行名 rownames(result) <- NULL
结果验证
运行上述任意一种方法后,result的结构和内容将与目标数据集data1完全一致。
内容的提问来源于stack exchange,提问作者Newcomer
相关产品推荐
相关产品推荐

