替换normal.meth行名时维度不匹配错误的技术求助
替换矩阵行名时维度不匹配报错解决方案
问题背景
现有两个无序矩阵normal.meth与meth.anno,需将normal.meth的行名替换为meth.anno中匹配probeID对应的geneNames取值,执行以下代码:
rownames(normal.meth) <- meth.anno[rownames(normal.meth) %in% meth.anno$probeID,]$geneNames
报错信息
Error in dimnames(x) <- dn : length of 'dimnames' [1] not equal to array extent
输入数据片段
normal.meth 前5行5列
> dput(normal.meth[1:5,1:5]) structure(c(0.515455282880639, 0.97045870602927, 0.901165700394357, 0.162692440621765, 0.89710904117382, 0.343674708139333, 0.972888059373319, 0.915128518830134, 0.138548281248091, 0.92038340114066, 0.489442046194044, 0.953176398928738, 0.9335837684512, 0.10877780962314, 0.908916456368568, 0.578367183337139, 0.974767379272895, 0.913897835838841, 0.113012642325218, 0.903143353383053, 0.547287820003221, 0.956942802183765, 0.897469710684965, 0.146303108636092, 0.933600786692477), dim = c(5L, 5L), dimnames = list( c("cg00000029", "cg00000108", "cg00000109", "cg00000165", "cg00000236"), c("TCGA.A4.7288.11", "TCGA.BQ.5880.11", "TCGA.BQ.5879.11", "TCGA.BQ.5883.11", "TCGA.BQ.7055.11")))
meth.anno 前5行
> dput(meth.anno[1:5,]) structure(list(CpG_chrm = c("chr1", "chr1", "chr1", "chr1", "chr1" ), CpG_beg = c(15864L, 18826L, 29406L, 29424L, 29434L), CpG_end = c(15866L, 18828L, 29408L, 29426L, 29436L), probe_strand = c("-", "-", "-", "-", "-"), probeID = c("cg00000109", "cg00000165", "cg12045430", "cg20826792", "cg00381604"), genesUniq = c("WASH7P", "MIR6859-1;WASH7P", "MIR1302-2;MIR1302-2HG;WASH7P", "MIR1302-2;MIR1302-2HG;WASH7P", "MIR1302-2;MIR1302-2HG;WASH7P"), geneNames = c("WASH7P", "WASH7P", "WASH7P", "WASH7P", "WASH7P"), transcriptTypes = c("unprocessed_pseudogene", "miRNA;unprocessed_pseudogene", "miRNA;lncRNA;lncRNA;unprocessed_pseudogene", "miRNA;lncRNA;lncRNA;unprocessed_pseudogene", "miRNA;lncRNA;lncRNA;unprocessed_pseudogene" ), transcriptIDs = c("ENST00000488147.1", "ENST00000619216.1;ENST00000488147.1", "ENST00000607096.1;ENST00000469289.1;ENST00000473358.1;ENST00000488147.1", "ENST00000607096.1;ENST00000469289.1;ENST00000473358.1;ENST00000488147.1", "ENST00000607096.1;ENST00000469289.1;ENST00000473358.1;ENST00000488147.1" ), distToTSS = c("13706", "-1390;10744", "-959;-860;-147;164", "-941;-842;-129;146", "-931;-832;-119;136"), CGI = c(NA, NA, "CGI:chr1:28735-29737", "CGI:chr1:28735-29737", "CGI:chr1:28735-29737" ), CGIposition = c(NA, NA, "Island", "Island", "Island")), row.names = c(NA, 5L), class = "data.frame")
维度信息
> dim(meth.anno) [1] 423686 12 > dim(normal.meth) [1] 486427 45
错误原因
原代码使用%in%筛选meth.anno时,仅返回meth.anno中存在于normal.meth行名的记录,但:
- 筛选结果的顺序与
normal.meth的行名顺序完全无关,无法对应到正确的行 - 当
normal.meth中存在probeID未出现在meth.anno中时,筛选结果的长度会小于normal.meth的行数,导致行名赋值时维度不匹配
解决方案
使用match()函数,该函数会返回normal.meth每个行名在meth.anno$probeID中的位置,无匹配项时返回NA,结果长度与normal.meth行数完全一致:
基础版本(无匹配时行名为NA)
rownames(normal.meth) <- meth.anno$geneNames[match(rownames(normal.meth), meth.anno$probeID)]
优化版本(无匹配时保留原probeID)
如果希望未匹配到的行保留原probeID作为行名,可添加判断:
# 先获取匹配的geneNames,无匹配则为NA gene_names <- meth.anno$geneNames[match(rownames(normal.meth), meth.anno$probeID)] # 替换行名,NA时保留原probeID rownames(normal.meth) <- ifelse(is.na(gene_names), rownames(normal.meth), gene_names)
内容的提问来源于stack exchange,提问作者Anon
相关产品推荐
相关产品推荐

