如何拆分数据框列中多分隔符字符串并提取去重ID前缀?
问题描述
原始数据框:
| Fasta headers |
|---|
| ab12_P002;ab12_P003;ab12_P005;ab23_P002;ab23_P001 |
| ab45_P001;ab36_P001 |
| ab55_P001;ab55_P002 |
已使用以下代码将分号分隔的字符串拆分为行:
library(tidyr) library(dplyr) without_02473 %>% mutate(`Fasta headers` = strsplit(as.character(`Fasta headers`), ";")) %>% unnest(`Fasta headers`)
得到中间结果:
| Fasta headers |
|---|
| ab12_P002 |
| ab12_P003 |
| ab12_P005 |
| ab23_P002 |
| ab23_P001 |
| ab45_P001 |
最终需要提取每个条目中_之前的前缀,去重后得到如下结果(排除ab55):
| Fasta headers |
|---|
| ab12 |
| ab23 |
| ab45 |
| ab36 |
尝试过分组、过滤等方法未成功,寻求解决办法。
解决方案
可以通过拆分提取前缀+去重+过滤的步骤实现需求,以下提供两种简洁的实现方式:
方法一:使用separate_rows+separate
library(tidyr) library(dplyr) without_02473 %>% # 按分号拆分字符串为独立行 separate_rows(`Fasta headers`, sep = ";") %>% # 按下划线拆分,提取前缀作为新的Fasta headers列 separate(`Fasta headers`, into = c(`Fasta headers`, "suffix"), sep = "_", remove = TRUE) %>% # 去除重复的前缀 distinct(`Fasta headers`) %>% # 过滤掉不需要的ab55条目 filter(`Fasta headers` != "ab55")
方法二:使用separate_rows+str_extract
library(tidyr) library(dplyr) library(stringr) without_02473 %>% separate_rows(`Fasta headers`, sep = ";") %>% # 提取下划线之前的所有字符作为前缀 mutate(`Fasta headers` = str_extract(`Fasta headers`, "^[^_]+")) %>% distinct(`Fasta headers`) %>% filter(`Fasta headers` != "ab55")
关键说明:
separate_rows是tidyr中专门用于拆分字符串为行的函数,比strsplit+unnest更简洁。str_extract(Fasta headers, "^[^_]+")的逻辑是:匹配从字符串开头到第一个下划线之前的所有字符,精准提取前缀。distinct用于去除重复的前缀,filter用于排除不需要的ab55(如果你的需求中不需要排除,可去掉这一行)。
内容的提问来源于stack exchange,提问作者code_newbie
相关产品推荐
相关产品推荐

