如何在R中基于向量为数据框的每行选择不同列
按行动态选择数据框列的高效实现
我有一个4列的数据框,需要为每行提取其中2列(每行选择的列不同)。
示例数据
repro = structure(list(c1 = c(0L, 0L, 1L, 1L, 0L, 1L), c2 = c(1L, 1L, 0L, 0L, 1L, 1L), c1 = c(0L, 1L, 1L, 0L, 1L, 0L), c2 = c(0L, 1L, 1L, 1L, 1L, 0L)), row.names = c(86L, 59L, 58L, 79L, 70L, 83L), class = "data.frame") head(repro) # c1 c2 c1 c2 # 86 0 1 0 0 # 59 0 1 1 1 # 58 1 0 1 1 # 79 1 0 0 1 # 70 0 1 1 1 # 83 1 1 0 0
列选择向量
用于指定每行要选择的两列的索引:
col.sel1 = c(2, 1, 2, 2, 2, 2) col.sel2 = c(4, 3, 3, 4, 3, 3)
低效的循环方法
我用循环实现了需求,但原始数据有数千行,运行速度极慢:
# 创建结果表 offspring = NULL for (i in 1:nrow(repro)) { offs = cbind(c3 = repro[i, col.sel1[i]], c4 = repro[i, col.sel2[i]]) offspring = rbind(offspring, offs) } head(offspring) # c3 c4 # [1,] 1 0 # [2,] 0 1 # [3,] 0 1 # [4,] 0 1 # [5,] 1 1 # [6,] 1 0
问题
有没有更快的方法,基于col.sel1和col.sel2这两个向量为每行选择不同的列?
我试过以下代码,但没得到预期结果:
rp[1:6, cs1] lapply(cs1, function(x) rp[,x])
高效解决方案
方法1:矩阵索引(基础R最快实现)
直接通过行号+列号组成的索引矩阵提取元素,这是基础R中性能最优的写法:
# 构造每行对应的索引对 idx1 = cbind(1:nrow(repro), col.sel1) idx2 = cbind(1:nrow(repro), col.sel2) # 提取元素并组合成数据框 offspring = data.frame( c3 = repro[idx1], c4 = repro[idx2] ) head(offspring) # c3 c4 # 1 1 0 # 2 0 1 # 3 0 1 # 4 0 1 # 5 1 1 # 6 1 0
方法2:tidyverse风格实现
如果习惯使用dplyr,可以结合rowwise()和pick()实现:
library(dplyr) offspring = repro %>% rowwise() %>% mutate( c3 = pick(all_of(col.sel1[[cur_row()]])), c4 = pick(all_of(col.sel2[[cur_row()]])) ) %>% ungroup() %>% select(c3, c4) head(offspring) # # A tibble: 6 × 2 # c3 c4 # <int> <int> # 1 1 0 # 2 0 1 # 3 0 1 # 4 0 1 # 5 1 1 # 6 1 0
方法3:向量化mapply操作
利用mapply逐行匹配行号和列号提取元素:
offspring = data.frame( c3 = mapply(function(row, col) repro[row, col], 1:nrow(repro), col.sel1), c4 = mapply(function(row, col) repro[row, col], 1:nrow(repro), col.sel2) ) head(offspring) # c3 c4 # 1 1 0 # 2 0 1 # 3 0 1 # 4 0 1 # 5 1 1 # 6 1 0
内容的提问来源于stack exchange,提问作者M. Beausoleil
相关产品推荐
相关产品推荐

