如何在R中将基因集因子表转换为基因存在-缺失(0/1)计数表?
解决GSEA分析中基因集-基因的0/1矩阵转换问题
问题场景
使用clusterProfiler进行GSEA分析时,提取leading edge相关数据后得到如下格式的数据集:
# 示例数据 GS1 <- c("a", "b", "c", "d", "e", "f") GS2 <- c("b", "c", "d", "e", "f", "g") GS3 <- c("a", "b", "c", NA,NA,NA) GS4 <- c("a", "d", "e", "g", NA, NA) GS5 <- c("a", "b", "c", "d", NA, NA) df <- data.frame(rbind(GS1, GS2, GS3, GS4, GS5))
当前每行对应一个基因集,列存储该基因集的基因(含NA),需要转换为**每行对应基因集、每列对应单个基因,值为1(基因存在于基因集)或0(不存在)**的矩阵,避免手动处理。
方法1:tidyverse工具链实现(推荐,代码简洁高效)
通过长-宽格式转换实现,适合大规模数据处理:
library(tidyverse) # 1. 为数据添加基因集名称列 df_with_gs <- df %>% mutate(gene_set = rownames(.)) # 2. 转换为长格式并过滤NA值 long_df <- df_with_gs %>% pivot_longer(cols = -gene_set, names_to = "temp_col", values_to = "gene") %>% filter(!is.na(gene)) %>% select(-temp_col) # 3. 转换为0/1矩阵格式 result_matrix <- long_df %>% mutate(value = 1) %>% pivot_wider(names_from = gene, values_from = value, values_fill = 0) %>% column_to_rownames("gene_set") # 查看结果 print(result_matrix)
方法2:base R原生实现(无需额外包)
通过提取唯一基因、逐行判断基因存在性实现:
# 1. 获取所有非NA的唯一基因列表 all_genes <- unique(unlist(df)) all_genes <- all_genes[!is.na(all_genes)] # 2. 逐行生成基因集的0/1向量 result_matrix <- t(apply(df, 1, function(row) { genes_in_set <- row[!is.na(row)] as.integer(all_genes %in% genes_in_set) })) # 设置行列名称 rownames(result_matrix) <- rownames(df) colnames(result_matrix) <- all_genes # 查看结果 print(result_matrix)
内容的提问来源于stack exchange,提问作者B_slash_
相关产品推荐
相关产品推荐

