在R语言中获取各模块颜色分组的行索引并生成列表
获取R数据框中不同分组的行索引列表
嘿,我来帮你搞定这个需求!首先咱们先把你提供的结构转换成可用的R数据框:
df <- structure(list(Gene_ID = structure(c(1L, 3L, 4L, 5L, 6L, 7L, 8L, 9L, 10L, 2L), .Label = c("g1", "g10", "g2", "g3", "g4", "g5", "g6", "g7", "g8", "g9"), class = "factor"), Module_Color = structure(c(3L, 1L, 3L, 2L, 3L, 1L, 2L, 3L, 2L, 1L), .Label = c("blue", "green", "red"), class = "factor")), .Names = c("Gene_ID", "Module_Color"), class = "data.frame", row.names = c(NA, -10L))
接下来给你三种实用的方法,按需选择就行:
方法1:基础R的split()函数(最简洁)
这是基础R里最直接的解决方案,一行代码就能搞定:
index_list <- split(seq_len(nrow(df)), df$Module_Color)
seq_len(nrow(df))生成从1到数据框行数的行索引,split()会自动按Module_Color的不同取值把这些索引拆分成列表,列表的名称就是对应的颜色分组。比如你想查看红色组的行索引,直接调用index_list$red就能得到c(1,3,5,8)。
方法2:基础R的tapply()函数
如果你更习惯用tapply,也能达到完全一样的效果:
index_list <- tapply(seq_len(nrow(df)), df$Module_Color, c)
这里tapply会对每个颜色分组的行索引应用c函数(其实就是把索引合并成向量),最终输出的列表和split方法的结果完全一致。
方法3:tidyverse工具链(适合tidy语法爱好者)
如果你平时常用dplyr这类tidyverse包,可以用这套流程:
library(dplyr) library(purrr) index_list <- df %>% mutate(row_index = row_number()) %>% # 先给每行添上行号 group_split(Module_Color, .keep = FALSE) %>% # 按颜色分组拆分数据框 set_names(unique(df$Module_Color)) %>% # 给列表加上颜色名称 map(pull, row_index) # 从每个子数据框里提取行索引
不管用哪种方法,最终得到的index_list都是这样的结构:
$blue [1] 2 6 10 $green [1] 4 7 9 $red [1] 1 3 5 8
内容的提问来源于stack exchange,提问作者J. Smith
相关产品推荐
相关产品推荐

