如何将GO_term-基因-Log2FC矩阵转换为带Log2FC的二元矩阵
转换GO术语-基因矩阵为二元关联矩阵(含Log2FC)
这里有个用R语言实现的方案,刚好能把你的三列长格式数据转换成目标的宽格式二元矩阵,还保留了Log2FC信息:
步骤1:准备输入数据
首先把你的原始数据转换成R数据框,如果是从文件读取的话,可以用read.table()或者read_csv(),这里先直接构造示例数据:
# 构造输入数据框 input_df <- data.frame( GO_term = c("cell adhesion", "cell adhesion", "cell adhesion", "cell-matrix adhesion", "cell-matrix adhesion", "positive regulation of cell migration", "positive regulation of cell migration", "cellular oxidant detoxification", "cellular oxidant detoxification", "muscle contraction", "muscle contraction"), Gene_Name = c("IGFBP7", "PVRL4", "NCAM1", "ITGA7", "ITGA4", "ITGA5", "RRAS2", "FABP1", "LTC4S", "ACTA2", "VCL"), Log2FC = c(1.38, -1.40, -1.35, -1.20, 0.75, -1.36, -0.59, 2.35, -0.59, -1.21, -1.06), stringsAsFactors = FALSE # 确保字符型不转成因子 )
步骤2:用tidyverse转换为宽格式二元矩阵
我们用tidyverse里的pivot_wider函数来完成转换,这个函数处理长转宽非常灵活:
# 加载tidyverse包(如果没安装先运行install.packages("tidyverse")) library(tidyverse) # 转换为目标格式 chord_df <- input_df %>% # 给每个基因-GO术语对标记为1(表示该基因属于这个GO术语) mutate(present = 1) %>% # 转换为宽格式:基因和Log2FC作为行标识,GO术语作为列,缺失值填充0 pivot_wider( id_cols = c(Gene_Name, Log2FC), names_from = GO_term, values_from = present, values_fill = 0 ) %>% # 把基因名设为行名(可选,方便后续做和弦图等可视化) column_to_rownames("Gene_Name")
步骤3:查看结果
运行完上面的代码后,chord_df就是你想要的矩阵了,查看前几行的话用head(chord_df),输出大概是这样:
Log2FC cell adhesion cell-matrix adhesion positive regulation of cell migration cellular oxidant detoxification muscle contraction IGFBP7 1.38 1 0 0 0 0 PVRL4 -1.40 1 0 0 0 0 NCAM1 -1.35 1 0 0 0 0 ITGA7 -1.20 0 1 0 0 0 ITGA4 0.75 0 1 0 0 0 ITGA5 -1.36 0 0 1 0 0
如果你习惯用reshape2包,也可以用dcast函数实现同样的效果:
# 加载reshape2(没安装先运行install.packages("reshape2")) library(reshape2) chord_df <- dcast(input_df, Gene_Name + Log2FC ~ GO_term, fun.aggregate = function(x) 1, fill = 0) rownames(chord_df) <- chord_df$Gene_Name chord_df$Gene_Name <- NULL
这个矩阵可以直接用来做和弦图(比如circlize包的chordDiagram()),完全匹配你给出的示例结构。
内容的提问来源于stack exchange,提问作者Zizogolu
相关产品推荐
相关产品推荐

