You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:如何在R中修改nwk树tip.labels,仅保留开头ID

在R中修改系统发育树的tip.label,仅保留编号ID

问题描述

我需要修改R里nwk格式系统发育树的tip.label,只保留每个标签开头的编号ID。比如当前标签是:
AB177299.1 Uncultured bacterium gene for 16S rRNA, clone: ODP1251B1.3
希望只留下AB177299.1。

一开始想过用gsub()处理tree$tip.label,但觉得每个标签的ID和后续文本都不一样,不知道怎么写正则。后来我用分类表的行名生成了包含所有ID的向量tip.ids,尝试用order()匹配替换,代码如下:

tip.ids <- rownames(taxonomy)
tree$tip.label[order(tree$tip.label) %in% tip.ids] <- tip.ids

但这段代码完全没效果,现在卡壳了。

复现数据

# tip.ids向量
tip.ids <- c("AB109878.1", "AB109879.1", "AB109880.1", "AB109881.1", 
             "AB109882.1", "AB109883.1", "AB109884.1", "AB109885.1")

# 系统发育树对象
tree <- structure(list(edge = structure(c(9L, 10L, 10L, 9L, 11L, 11L, 
12L, 13L, 13L, 12L, 14L, 14L, 15L, 15L, 10L, 1L, 2L, 11L, 3L, 
12L, 13L, 4L, 5L, 14L, 6L, 15L, 7L, 8L), dim = c(14L, 2L)), edge.length = c(0.0341921975, 
5e-09, 0.12821348, 0.000367458500000008, 0.027617765, 0.037677039, 
0.028633124, 0.014468092, 5e-09, 0.009763081, 0.078168769, 0.021640684, 
0.341568464, 0.092957415), Nnode = 7L, node.label = c("root", 
"0.917", "", "0.929", "0.921", "0.302", "0.692"), tip.label = c("'AB109881.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-4'", 
"'AB109880.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-3'", 
"'AB109883.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-6'", 
"'AB109879.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-2'", 
"'AB109884.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-7'", 
"'AB109878.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-1'", 
"'AB109882.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-5'", 
"'AB109885.1 Uncultured archaeon gene for 16S rRNA, partial sequence, clone:pMLA-8'"
)), class = "phylo", order = "cladewise")

解决方法

方法1:正则表达式快速提取(最省事)

其实gsub()完全能搞定,你的担心多余了——不管后续文本是什么,只要ID是标签开头、空格前的内容,就能用正则提取。另外注意你的tip.label里还带首尾单引号,得先去掉:

# 移除标签首尾的单引号
tree$tip.label <- gsub("^'|'$", "", tree$tip.label)
# 提取第一个空格前的所有内容(也就是目标ID)
tree$tip.label <- gsub("\\s.*", "", tree$tip.label)
  • ^'|'$:匹配开头或结尾的单引号,替换为空;
  • \\s.*:匹配第一个空格及后面的所有字符,替换为空,剩下的就是ID。

方法2:用tip.ids精准匹配替换

你之前的代码没生效,是因为order(tree$tip.label)返回的是排序后的索引,而tree$tip.label本身是带单引号和后续文本的完整字符串,根本不可能和tip.ids里的纯ID匹配,所以筛选出来的索引是空的,自然没修改。

可以用grepl()匹配每个标签里的ID,再替换:

# 先移除单引号
tree$tip.label <- gsub("^'|'$", "", tree$tip.label)
# 遍历每个ID,找到对应标签并替换
for (id in tip.ids) {
  tree$tip.label[grepl(paste0("^", id), tree$tip.label)] <- id
}

或者用sapply()简化写法:

tree$tip.label <- gsub("^'|'$", "", tree$tip.label)
tree$tip.label <- sapply(tree$tip.label, function(x) {
  # 找到当前标签开头匹配的ID
  matched_id <- tip.ids[grepl(paste0("^", tip.ids), x)]
  # 有匹配就用ID,没匹配就保留原内容
  if (length(matched_id) > 0) matched_id else x
})

内容的提问来源于stack exchange,提问作者Geomicro

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.11 00:47:32