You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

按probeID-target_symbol对出现次数重复meth/exp行的R代码报错排查

问题解决:根据配对出现次数重复矩阵行

需求背景

meth的行名与distal.dmr.cpgs的probeID列匹配,exp的行名与distal.dmr.cpgs的target_symbol列匹配。需要根据distal.dmr.cpgs中同一probeID和target_symbol配对的出现次数,重复meth和exp对应的行。

尝试代码

# 按`probeID`和`target_symbol`配对的出现次数,重复`meth`和`exp`的行
# `probeID`匹配`meth`的行名;`target_symbol`匹配`exp`的行名
n.occur <- distal.dmr.cpgs %>% count(probeID, target_symbol)
if(n.occur$probeID == rownames(meth)) {
  for (n in n.occur$n){
    meth <- meth[rep(seq_len(nrow(meth)), each=n), ]
    exp <- exp[rep(seq_len(nrow(exp)), each=n), ]
  }
}

报错信息

Error in if (n.occur$probeID == rownames(meth)) { : 
  the condition has length > 1

输入数据

> dput(meth[1:5,1:5])
structure(c(0.965846994535519, 0.574436129165911, 0.8722836745901, 
0.101103884990587, 0.408252566147303, 0.945927883796075, 0.480092866016333, 
0.853431761902537, 0.0755436985006079, 0.448047437384109, 0.95713326619671, 
0.8891500945397, 0.926905435737289, 0.203369962221719, 0.644251635185525, 
0.961321856037767, 0.501968996835947, 0.865640773821373, 0.11543134872418, 
0.472688968983197, 0.941828841874757, 0.484359313077939, 0.9011216776396, 
0.155463949005687, 0.566371097999143), dim = c(5L, 5L), dimnames = list(
    c("cg03477043", "cg00926926", "cg00488747", "cg04452095", 
    "cg04658243"), c("TCGA.Y8.A8S1.01", "TCGA.Y8.A8S0.01", "TCGA.Y8.A8RZ.01", 
    "TCGA.Y8.A8RY.01", "TCGA.Y8.A897.01")))

> dput(exp[1:5,1:5])
structure(c(8.8764930371619, 7.10418439636105, 5.82248600796462, 
9.6336088827523, 7.12980348328146, 8.50488041264077, 7.4053080406626, 
6.10447768176044, 9.82618201799457, 7.25526533976169, 8.91383556515162, 
7.69822939472376, 6.293732617918, 10.0591928112391, 7.25510044887476, 
8.05139916296298, 7.19502387866043, 6.05119750390752, 9.29733149115023, 
7.55517805926009, 9.13478579142527, 7.38961451449646, 6.09188514895038, 
9.67121455892562, 7.28978856899351), dim = c(5L, 5L), dimnames = list(
    c("SRRM2", "ANKRD11", "RPTOR", "HSP90AA1", "RER1"), c("TCGA.2K.A9WE.01", 
    "TCGA.2Z.A9J1.01", "TCGA.2Z.A9J3.01", "TCGA.2Z.A9J5.01", 
    "TCGA.2Z.A9J6.01")))

> dput(distal.dmr.cpgs[,c("probeID","target_symbol")][1:5,])
structure(list(probeID = c("cg03477043", "cg03477043", "cg00926926", 
"cg00926926", "cg00926926"), target_symbol = c("SRRM2", "SRRM2", 
"ANKRD11", "ANKRD11", "ANKRD11")), row.names = c(785L, 786L, 
866L, 867L, 868L), class = "data.frame")

预期输出

cg03477043和SRRM2对应的行需在meth和exp中重复2次,cg00926926和ANKRD11对应的行需重复3次。


解决方案

错误原因

报错的核心是if(n.occur$probeID == rownames(meth))这行代码:n.occur$probeID和rownames(meth)都是向量,直接用==会得到一个布尔向量,而if语句只能接受长度为1的条件,因此触发错误。另外原代码的循环逻辑错误,会把所有行重复多次,而非按配对的次数单独处理对应行。

修正代码

我们可以通过生成对应行的重复索引来实现需求,以下是简洁的实现方式:

library(dplyr)

# 统计每个配对的出现次数
n.occur <- distal.dmr.cpgs %>% count(probeID, target_symbol)

# 处理meth:先匹配probeID对应的行号,再按次数重复
meth_rows <- match(n.occur$probeID, rownames(meth))
meth_repeated <- meth[rep(meth_rows, n.occur$n), ]

# 处理exp:先匹配target_symbol对应的行号,再按次数重复
exp_rows <- match(n.occur$target_symbol, rownames(exp))
exp_repeated <- exp[rep(exp_rows, n.occur$n), ]

结果验证

运行上述代码后,meth_repeated中cg03477043会出现2次,cg00926926出现3次;exp_repeated中SRRM2出现2次,ANKRD11出现3次,完全符合预期输出。


内容的提问来源于stack exchange,提问作者Anon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 00:56:59