求助:如何在R中修改nwk树tip.labels,仅保留开头ID
在R中修改系统发育树的tip.label,仅保留编号ID
问题描述
我需要修改R里nwk格式系统发育树的tip.label,只保留每个标签开头的编号ID。比如当前标签是:AB177299.1 Uncultured bacterium gene for 16S rRNA, clone: ODP1251B1.3
希望只留下AB177299.1。
一开始想过用gsub()处理tree$tip.label,但觉得每个标签的ID和后续文本都不一样,不知道怎么写正则。后来我用分类表的行名生成了包含所有ID的向量tip.ids,尝试用order()匹配替换,代码如下:
tip.ids <- rownames(taxonomy) tree$tip.label[order(tree$tip.label) %in% tip.ids] <- tip.ids
但这段代码完全没效果,现在卡壳了。
复现数据
# tip.ids向量 tip.ids <- c("AB109878.1", "AB109879.1", "AB109880.1", "AB109881.1", "AB109882.1", "AB109883.1", "AB109884.1", "AB109885.1") # 系统发育树对象 tree <- structure(list(edge = structure(c(9L, 10L, 10L, 9L, 11L, 11L, 12L, 13L, 13L, 12L, 14L, 14L, 15L, 15L, 10L, 1L, 2L, 11L, 3L, 12L, 13L, 4L, 5L, 14L, 6L, 15L, 7L, 8L), dim = c(14L, 2L)), edge.length = c(0.0341921975, 5e-09, 0.12821348, 0.000367458500000008, 0.027617765, 0.037677039, 0.028633124, 0.014468092, 5e-09, 0.009763081, 0.078168769, 0.021640684, 0.341568464, 0.092957415), Nnode = 7L, node.label = c("root", "0.917", "", "0.929", "0.921", "0.302", "0.692"), tip.label = c("'AB109881.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-4'", "'AB109880.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-3'", "'AB109883.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-6'", "'AB109879.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-2'", "'AB109884.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-7'", "'AB109878.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-1'", "'AB109882.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-5'", "'AB109885.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-8'" )), class = "phylo", order = "cladewise")
解决方法
方法1:正则表达式快速提取(最省事)
其实gsub()完全能搞定,你的担心多余了——不管后续文本是什么,只要ID是标签开头、空格前的内容,就能用正则提取。另外注意你的tip.label里还带首尾单引号,得先去掉:
# 移除标签首尾的单引号 tree$tip.label <- gsub("^'|'$", "", tree$tip.label) # 提取第一个空格前的所有内容(也就是目标ID) tree$tip.label <- gsub("\\s.*", "", tree$tip.label)
^'|'$:匹配开头或结尾的单引号,替换为空;\\s.*:匹配第一个空格及后面的所有字符,替换为空,剩下的就是ID。
方法2:用tip.ids精准匹配替换
你之前的代码没生效,是因为order(tree$tip.label)返回的是排序后的索引,而tree$tip.label本身是带单引号和后续文本的完整字符串,根本不可能和tip.ids里的纯ID匹配,所以筛选出来的索引是空的,自然没修改。
可以用grepl()匹配每个标签里的ID,再替换:
# 先移除单引号 tree$tip.label <- gsub("^'|'$", "", tree$tip.label) # 遍历每个ID,找到对应标签并替换 for (id in tip.ids) { tree$tip.label[grepl(paste0("^", id), tree$tip.label)] <- id }
或者用sapply()简化写法:
tree$tip.label <- gsub("^'|'$", "", tree$tip.label) tree$tip.label <- sapply(tree$tip.label, function(x) { # 找到当前标签开头匹配的ID matched_id <- tip.ids[grepl(paste0("^", tip.ids), x)] # 有匹配就用ID,没匹配就保留原内容 if (length(matched_id) > 0) matched_id else x })
内容的提问来源于stack exchange,提问作者Geomicro
相关产品推荐
相关产品推荐

