按probeID-target_symbol对出现次数重复meth/exp行的R代码报错排查
问题解决:根据配对出现次数重复矩阵行
需求背景
meth的行名与distal.dmr.cpgs的probeID列匹配,exp的行名与distal.dmr.cpgs的target_symbol列匹配。需要根据distal.dmr.cpgs中同一probeID和target_symbol配对的出现次数,重复meth和exp对应的行。
尝试代码
# 按`probeID`和`target_symbol`配对的出现次数,重复`meth`和`exp`的行 # `probeID`匹配`meth`的行名;`target_symbol`匹配`exp`的行名 n.occur <- distal.dmr.cpgs %>% count(probeID, target_symbol) if(n.occur$probeID == rownames(meth)) { for (n in n.occur$n){ meth <- meth[rep(seq_len(nrow(meth)), each=n), ] exp <- exp[rep(seq_len(nrow(exp)), each=n), ] } }
报错信息
Error in if (n.occur$probeID == rownames(meth)) { : the condition has length > 1
输入数据
> dput(meth[1:5,1:5]) structure(c(0.965846994535519, 0.574436129165911, 0.8722836745901, 0.101103884990587, 0.408252566147303, 0.945927883796075, 0.480092866016333, 0.853431761902537, 0.0755436985006079, 0.448047437384109, 0.95713326619671, 0.8891500945397, 0.926905435737289, 0.203369962221719, 0.644251635185525, 0.961321856037767, 0.501968996835947, 0.865640773821373, 0.11543134872418, 0.472688968983197, 0.941828841874757, 0.484359313077939, 0.9011216776396, 0.155463949005687, 0.566371097999143), dim = c(5L, 5L), dimnames = list( c("cg03477043", "cg00926926", "cg00488747", "cg04452095", "cg04658243"), c("TCGA.Y8.A8S1.01", "TCGA.Y8.A8S0.01", "TCGA.Y8.A8RZ.01", "TCGA.Y8.A8RY.01", "TCGA.Y8.A897.01"))) > dput(exp[1:5,1:5]) structure(c(8.8764930371619, 7.10418439636105, 5.82248600796462, 9.6336088827523, 7.12980348328146, 8.50488041264077, 7.4053080406626, 6.10447768176044, 9.82618201799457, 7.25526533976169, 8.91383556515162, 7.69822939472376, 6.293732617918, 10.0591928112391, 7.25510044887476, 8.05139916296298, 7.19502387866043, 6.05119750390752, 9.29733149115023, 7.55517805926009, 9.13478579142527, 7.38961451449646, 6.09188514895038, 9.67121455892562, 7.28978856899351), dim = c(5L, 5L), dimnames = list( c("SRRM2", "ANKRD11", "RPTOR", "HSP90AA1", "RER1"), c("TCGA.2K.A9WE.01", "TCGA.2Z.A9J1.01", "TCGA.2Z.A9J3.01", "TCGA.2Z.A9J5.01", "TCGA.2Z.A9J6.01"))) > dput(distal.dmr.cpgs[,c("probeID","target_symbol")][1:5,]) structure(list(probeID = c("cg03477043", "cg03477043", "cg00926926", "cg00926926", "cg00926926"), target_symbol = c("SRRM2", "SRRM2", "ANKRD11", "ANKRD11", "ANKRD11")), row.names = c(785L, 786L, 866L, 867L, 868L), class = "data.frame")
预期输出
cg03477043和SRRM2对应的行需在meth和exp中重复2次,cg00926926和ANKRD11对应的行需重复3次。
解决方案
错误原因
报错的核心是if(n.occur$probeID == rownames(meth))这行代码:n.occur$probeID和rownames(meth)都是向量,直接用==会得到一个布尔向量,而if语句只能接受长度为1的条件,因此触发错误。另外原代码的循环逻辑错误,会把所有行重复多次,而非按配对的次数单独处理对应行。
修正代码
我们可以通过生成对应行的重复索引来实现需求,以下是简洁的实现方式:
library(dplyr) # 统计每个配对的出现次数 n.occur <- distal.dmr.cpgs %>% count(probeID, target_symbol) # 处理meth:先匹配probeID对应的行号,再按次数重复 meth_rows <- match(n.occur$probeID, rownames(meth)) meth_repeated <- meth[rep(meth_rows, n.occur$n), ] # 处理exp:先匹配target_symbol对应的行号,再按次数重复 exp_rows <- match(n.occur$target_symbol, rownames(exp)) exp_repeated <- exp[rep(exp_rows, n.occur$n), ]
结果验证
运行上述代码后,meth_repeated中cg03477043会出现2次,cg00926926出现3次;exp_repeated中SRRM2出现2次,ANKRD11出现3次,完全符合预期输出。
内容的提问来源于stack exchange,提问作者Anon
相关产品推荐
相关产品推荐

