按组随机选取同一行的value与treatment并替换组内对应列
问题描述
现有如下R数据框:
data <- structure(list(group = c(1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3), value = c("1", "4", "5", "3", "3", "5", "3", "5", "1", "4", "2"), treatment = c("A", "A", "B", "C", "C", "C", "C", "C", "A", "A","NA"), retain = c("88", "664", "797", "131", "88", "997", "645", "88", "13", "79", "5")), class = "data.frame", row.names = c(NA, -11L))
需要对每个group执行以下操作:
- 随机选取该组内某一行的
value和treatment值 - 将该组内所有行的
value和treatment列替换为选中的值 - 保留
retain列的原有数据
期望输出示例(随机选取结果仅作参考,每次运行结果会不同):
data <- structure(list(group = c(1, 1, 1, 1, 2, 2, 2, 2, 3, 3, 3), value = c("4", "4", "4", "4", "5", "5", "5", "5", "2", "2", "2"), treatment = c("A", "A", "A", "A", "C", "C", "C", "C", "NA", "NA","NA"), retain = c("88", "664", "797", "131", "88", "997", "645", "88", "13", "79", "5")), class = "data.frame", row.names = c(NA, -11L))
解决方案
方法1:使用dplyr包(推荐,代码简洁易读)
先确保已安装dplyr包,未安装则运行install.packages("dplyr"):
library(dplyr) set.seed(123) # 设置随机种子,确保结果可重复,不需要可移除 result <- data %>% group_by(group) %>% mutate( sample_idx = sample(n(), 1), # 随机生成组内一行的索引 value = first(value[sample_idx]), # 用选中行的value替换整组 treatment = first(treatment[sample_idx]) # 用选中行的treatment替换整组 ) %>% select(-sample_idx) %>% # 删除临时生成的索引列 ungroup() # 取消分组,返回普通数据框 # 查看结果 print(result)
方法2:Base R实现(无需额外包)
set.seed(123) # 按group拆分数据框 split_data <- split(data, data$group) # 逐个处理分组 processed_data <- lapply(split_data, function(df) { # 随机选取组内一行 sample_row <- df[sample(nrow(df), 1), ] # 替换整组的value和treatment df$value <- sample_row$value df$treatment <- sample_row$treatment return(df) }) # 合并回完整数据框并重置行名 result <- do.call(rbind, processed_data) row.names(result) <- NULL # 查看结果 print(result)
内容的提问来源于stack exchange,提问作者kalex
相关产品推荐
相关产品推荐

