You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拆分数据框列中多分隔符字符串并提取去重ID前缀?

问题描述

原始数据框:

Fasta headers
ab12_P002;ab12_P003;ab12_P005;ab23_P002;ab23_P001
ab45_P001;ab36_P001
ab55_P001;ab55_P002

已使用以下代码将分号分隔的字符串拆分为行:

library(tidyr)
library(dplyr)
without_02473 %>% 
  mutate(`Fasta headers` = strsplit(as.character(`Fasta headers`), ";")) %>%  
  unnest(`Fasta headers`) 

得到中间结果:

Fasta headers
ab12_P002
ab12_P003
ab12_P005
ab23_P002
ab23_P001
ab45_P001

最终需要提取每个条目中_之前的前缀,去重后得到如下结果(排除ab55):

Fasta headers
ab12
ab23
ab45
ab36

尝试过分组、过滤等方法未成功,寻求解决办法。


解决方案

可以通过拆分提取前缀+去重+过滤的步骤实现需求,以下提供两种简洁的实现方式:

方法一:使用separate_rows+separate

library(tidyr)
library(dplyr)

without_02473 %>%
  # 按分号拆分字符串为独立行
  separate_rows(`Fasta headers`, sep = ";") %>%
  # 按下划线拆分,提取前缀作为新的Fasta headers列
  separate(`Fasta headers`, into = c(`Fasta headers`, "suffix"), sep = "_", remove = TRUE) %>%
  # 去除重复的前缀
  distinct(`Fasta headers`) %>%
  # 过滤掉不需要的ab55条目
  filter(`Fasta headers` != "ab55")

方法二:使用separate_rows+str_extract

library(tidyr)
library(dplyr)
library(stringr)

without_02473 %>%
  separate_rows(`Fasta headers`, sep = ";") %>%
  # 提取下划线之前的所有字符作为前缀
  mutate(`Fasta headers` = str_extract(`Fasta headers`, "^[^_]+")) %>%
  distinct(`Fasta headers`) %>%
  filter(`Fasta headers` != "ab55")

关键说明:

  • separate_rows是tidyr中专门用于拆分字符串为行的函数,比strsplit+unnest更简洁。
  • str_extract(Fasta headers, "^[^_]+")的逻辑是:匹配从字符串开头到第一个下划线之前的所有字符,精准提取前缀。
  • distinct用于去除重复的前缀,filter用于排除不需要的ab55(如果你的需求中不需要排除,可去掉这一行)。

内容的提问来源于stack exchange,提问作者code_newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 16:15:50