R语言dataframe按条件随机选中study并删除指定outcome行的方法
R实现方案
这个需求可以在R中轻松实现,以下是具体实现代码和说明:
首先是你提供的示例数据构造代码:
m = " study group outcome 1 1 1 A 2 1 1 B 3 1 2 A 4 1 2 B 5 2 1 A 6 2 1 B 7 2 2 A 8 2 2 B 9 3 1 B 10 4 1 B " data <- read.table(text=m,h=T)
基于dplyr的实现(代码更简洁易读)
library(dplyr) # 自定义参数,可根据需求修改 target_outcome <- "A" # 需要删除的outcome指定值 n <- 1 # 随机选取的study数量 # 步骤1:筛选出outcome取值数量大于1的合格study eligible_studies <- data %>% group_by(study) %>% filter(n_distinct(outcome) > 1) %>% pull(study) %>% unique() # 步骤2:从合格study中随机抽取n个 set.seed(123) # 可选,添加后可以复现随机抽样结果,不需要可删除 selected_studies <- sample(eligible_studies, size = n) # 步骤3:删除选中study内outcome等于指定值的行 result <- data %>% filter(!(study %in% selected_studies & outcome == target_outcome))
原生base R实现(无需安装第三方包)
# 自定义参数 target_outcome <- "A" n <- 1 # 步骤1:筛选合格study outcome_count <- tapply(data$outcome, data$study, function(x) length(unique(x))) eligible_studies <- as.numeric(names(outcome_count)[outcome_count > 1]) # 步骤2:随机抽样 set.seed(123) selected_studies <- sample(eligible_studies, size = n) # 步骤3:过滤行得到结果 result <- data[!(data$study %in% selected_studies & data$outcome == target_outcome), ]
效果验证
以示例参数target_outcome = "A"、n=1为例,如果随机抽中study=1,输出的result中study=1的所有outcome为"A"的行都会被删除,仅保留outcome为"B"的行;study=2、3、4的数据保持不变。
内容的提问来源于stack exchange,提问作者Simon Harmel
相关产品推荐
相关产品推荐

