You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中将基因集因子表转换为基因存在-缺失(0/1)计数表?

解决GSEA分析中基因集-基因的0/1矩阵转换问题

问题场景

使用clusterProfiler进行GSEA分析时,提取leading edge相关数据后得到如下格式的数据集:

# 示例数据
GS1 <- c("a", "b", "c", "d", "e", "f") 
GS2 <- c("b", "c", "d", "e", "f", "g") 
GS3 <- c("a", "b", "c", NA,NA,NA) 
GS4 <- c("a", "d", "e", "g", NA, NA) 
GS5 <- c("a", "b", "c", "d", NA, NA) 
df <- data.frame(rbind(GS1, GS2, GS3, GS4, GS5))

当前每行对应一个基因集,列存储该基因集的基因(含NA),需要转换为**每行对应基因集、每列对应单个基因,值为1(基因存在于基因集)或0(不存在)**的矩阵,避免手动处理。


方法1:tidyverse工具链实现(推荐,代码简洁高效)

通过长-宽格式转换实现,适合大规模数据处理:

library(tidyverse)

# 1. 为数据添加基因集名称列
df_with_gs <- df %>% 
  mutate(gene_set = rownames(.))

# 2. 转换为长格式并过滤NA值
long_df <- df_with_gs %>% 
  pivot_longer(cols = -gene_set, names_to = "temp_col", values_to = "gene") %>% 
  filter(!is.na(gene)) %>% 
  select(-temp_col)

# 3. 转换为0/1矩阵格式
result_matrix <- long_df %>% 
  mutate(value = 1) %>% 
  pivot_wider(names_from = gene, values_from = value, values_fill = 0) %>% 
  column_to_rownames("gene_set")

# 查看结果
print(result_matrix)

方法2:base R原生实现(无需额外包)

通过提取唯一基因、逐行判断基因存在性实现:

# 1. 获取所有非NA的唯一基因列表
all_genes <- unique(unlist(df))
all_genes <- all_genes[!is.na(all_genes)]

# 2. 逐行生成基因集的0/1向量
result_matrix <- t(apply(df, 1, function(row) {
  genes_in_set <- row[!is.na(row)]
  as.integer(all_genes %in% genes_in_set)
}))

# 设置行列名称
rownames(result_matrix) <- rownames(df)
colnames(result_matrix) <- all_genes

# 查看结果
print(result_matrix)

内容的提问来源于stack exchange,提问作者B_slash_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 02:33:16