如何将TCGA肿瘤样本ID列拆分为多列?R语言实现遇问题
解决TCGA样本ID拆分问题
问题说明
现有一批TCGA肿瘤样本ID:
SAMPLE_ID TCGA.13.1407.01A.01R.1565.13 TCGA.24.2254.01A.01R.1568.13 TCGA.24.0982.01A.01R.1565.13 TCGA.24.1847.01A.01R.1566.13 TCGA.24.2289.01A.01R.1568.13 TCGA.31.1959.01A.01R.1568.13
需要将每个ID拆分为以下9个字段:Project、TSS、Participant、Sample、Vial、Portion、Analyte、Plate、Center,预期拆分示例:
SAMPLE_ID Project TSS Participant Sample Vial Portion Analyte Plate Center TCGA.13.1407.01A.01R.1565.13 TCGA 13 1407 01 A 01 R 1565 13
错误尝试及问题
原代码仅按点(.)拆分,导致01A、01R这类数字+字母的混合字段未被拆分为对应子字段,得到错误结果:
library(tidyr) library(dplyr) df = data.frame(SAMPLE_ID = c("TCGA.13.1407.01A.01R.1565.13", "TCGA.24.2254.01A.01R.1568.13", "TCGA.24.0982.01A.01R.1565.13", "TCGA.24.1847.01A.01R.1566.13", "TCGA.24.2289.01A.01R.1568.13", "TCGA.31.1959.01A.01R.1568.13")) result = df %>% separate(SAMPLE_ID, into = c("Project", "TSS", "Participant", "Sample", "Vial", "Portion", "Analyte", "Plate", "Center"), sep = "\\.")
错误输出:
Project TSS Participant Sample Vial Portion Analyte Plate Center <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> <chr> TCGA 13 1407 01A 01R 1565 13 NA NA TCGA 24 2254 01A 01R 1568 13 NA NA TCGA 24 0982 01A 01R 1565 13 NA NA TCGA 24 1847 01A 01R 1566 13 NA NA TCGA 24 2289 01A 01R 1568 13 NA NA TCGA 31 1959 01A 01R 1568 13 NA NA
解决方案
方法1:分步拆分(清晰易读)
先按点拆分出基础字段,再拆分数字+字母的混合字段:
library(tidyr) library(dplyr) df <- data.frame(SAMPLE_ID = c("TCGA.13.1407.01A.01R.1565.13", "TCGA.24.2254.01A.01R.1568.13", "TCGA.24.0982.01A.01R.1565.13", "TCGA.24.1847.01A.01R.1566.13", "TCGA.24.2289.01A.01R.1568.13", "TCGA.31.1959.01A.01R.1568.13")) # 1. 按点拆分出7个临时字段 temp_df <- df %>% separate(SAMPLE_ID, into = c("Project", "TSS", "Participant", "SampleVial", "PortionAnalyte", "Plate", "Center"), sep = "\\.") # 2. 拆分SampleVial为Sample(数字)和Vial(字母) temp_df <- temp_df %>% separate(SampleVial, into = c("Sample", "Vial"), sep = "(?<=\\d)(?=[A-Z])") # 3. 拆分PortionAnalyte为Portion(数字)和Analyte(字母) final_df <- temp_df %>% separate(PortionAnalyte, into = c("Portion", "Analyte"), sep = "(?<=\\d)(?=[A-Z])") # 补充原SAMPLE_ID列并调整顺序 final_df <- final_df %>% mutate(SAMPLE_ID = df$SAMPLE_ID) %>% select(SAMPLE_ID, Project, TSS, Participant, Sample, Vial, Portion, Analyte, Plate, Center)
方法2:一次性拆分(简洁高效)
使用正则表达式同时匹配点分隔和数字-字母分隔,一次性完成拆分:
library(tidyr) library(dplyr) df <- data.frame(SAMPLE_ID = c("TCGA.13.1407.01A.01R.1565.13", "TCGA.24.2254.01A.01R.1568.13", "TCGA.24.0982.01A.01R.1565.13", "TCGA.24.1847.01A.01R.1566.13", "TCGA.24.2289.01A.01R.1568.13", "TCGA.31.1959.01A.01R.1568.13")) final_df <- df %>% separate(SAMPLE_ID, into = c("Project", "TSS", "Participant", "Sample", "Vial", "Portion", "Analyte", "Plate", "Center"), sep = "\\.|(?<=\\d)(?=[A-Z])")
结果验证
最终输出符合预期:
SAMPLE_ID Project TSS Participant Sample Vial Portion Analyte Plate Center 1 TCGA.13.1407.01A.01R.1565.13 TCGA 13 1407 01 A 01 R 1565 13 2 TCGA.24.2254.01A.01R.1568.13 TCGA 24 2254 01 A 01 R 1568 13 3 TCGA.24.0982.01A.01R.1565.13 TCGA 24 0982 01 A 01 R 1565 13 4 TCGA.24.1847.01A.01R.1566.13 TCGA 24 1847 01 A 01 R 1566 13 5 TCGA.24.2289.01A.01R.1568.13 TCGA 24 2289 01 A 01 R 1568 13 6 TCGA.31.1959.01A.01R.1568.13 TCGA 31 1959 01 A 01 R 1568 13
内容的提问来源于stack exchange,提问作者nicholaspooran
相关产品推荐
相关产品推荐

