如何利用矩阵的index列作为索引修改矩阵值?
矩阵元素按历史index列替换为0
原始数据
我们有如下R矩阵:
M=structure(c(1, 2, 2, 2, 1, 3, 3, 4, 4, 6, 5, 5, 5, 6, 5, 5, 2, 2, 4, 1, 3, 1, 6, 4, 2, 5, 1, 3, 4, 6, 0.849542987113752, 0.849542987113752, 0.788371730579003, 0.788371730579003, 0.788371730579003, 0.788371730579003 ), .Dim = c(6L, 6L), .Dimnames = list(NULL, c("", "", "", "", "index", "cum_detection")))
该矩阵已按最后一列排序,显示如下:
index cum_detection [1,] 1 3 5 4 2 0.8495430 [2,] 2 4 6 1 5 0.8495430 [3,] 2 4 5 3 1 0.7883717 [4,] 2 6 5 1 3 0.7883717 [5,] 1 5 2 6 4 0.7883717 [6,] 3 5 2 4 6 0.7883717
期望输出
希望得到如下处理后的矩阵:
index cum_detection [1,] 1 3 5 4 2 0.8495430 [2,] 0 4 6 1 5 0.8495430 [3,] 0 4 0 3 1 0.7883717 [4,] 0 6 0 0 3 0.7883717 [5,] 0 0 0 6 4 0.7883717 [6,] 0 0 0 0 6 0.7883717
处理规则
对每行的第1至4列,若该列的值出现在之前所有已遍历行的index列中,则将其替换为0。
举例:第三行的第1-4列值为2 4 5 3,其中2和5已出现在前两行的index列(即M[1:2,"index"]的2和5),所以替换为0,得到0 4 0 3。
问题难点
尝试用apply函数处理,但无法在循环中获取当前行的索引,无法定位“之前已遍历行”的范围。
解决方案
方法1:基础R循环实现
通过for循环逐行处理,同时维护一个已出现的index集合,逻辑清晰易理解:
# 复制原始矩阵避免修改原数据 result <- M # 初始化已出现的index集合 seen_indexes <- c() for (i in 1:nrow(result)) { # 获取当前行的1-4列元素 current_cols <- result[i, 1:4] # 获取当前行的index值 current_index <- result[i, "index"] # 将1-4列中属于已出现index的元素替换为0 current_cols[current_cols %in% seen_indexes] <- 0 # 更新结果矩阵的1-4列 result[i, 1:4] <- current_cols # 将当前行的index加入已出现集合 seen_indexes <- c(seen_indexes, current_index) } # 查看处理后的结果 result
方法2:用purrr包的imap函数(带索引迭代)
如果习惯使用tidyverse工具链,可以用purrr::imap获取行索引,同时维护状态:
library(purrr) # 初始化已出现index集合 seen_indexes <- c() result <- imap_dfr(1:nrow(M), function(i, .) { row <- M[i, ] # 替换1-4列中符合条件的元素为0 row[1:4][row[1:4] %in% seen_indexes] <- 0 # 更新已出现的index集合 seen_indexes <<- c(seen_indexes, row["index"]) # 返回处理后的行 as.data.frame(t(row)) }) # 转换回矩阵格式 result <- as.matrix(result)
运行上述任意一种方法,均可得到期望的输出结果。
内容的提问来源于stack exchange,提问作者Tou Mou
相关产品推荐
相关产品推荐

