You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将TCGA肿瘤样本ID列拆分为多列?R语言实现遇问题

解决TCGA样本ID拆分问题

问题说明

现有一批TCGA肿瘤样本ID:

SAMPLE_ID
TCGA.13.1407.01A.01R.1565.13
TCGA.24.2254.01A.01R.1568.13
TCGA.24.0982.01A.01R.1565.13
TCGA.24.1847.01A.01R.1566.13
TCGA.24.2289.01A.01R.1568.13
TCGA.31.1959.01A.01R.1568.13

需要将每个ID拆分为以下9个字段:Project、TSS、Participant、Sample、Vial、Portion、Analyte、Plate、Center,预期拆分示例:

SAMPLE_ID                       Project TSS Participant Sample  Vial    Portion Analyte Plate   Center
TCGA.13.1407.01A.01R.1565.13    TCGA    13  1407        01      A       01      R       1565    13

错误尝试及问题

原代码仅按点(.)拆分,导致01A、01R这类数字+字母的混合字段未被拆分为对应子字段,得到错误结果:

library(tidyr)
library(dplyr)
df = data.frame(SAMPLE_ID = c("TCGA.13.1407.01A.01R.1565.13", "TCGA.24.2254.01A.01R.1568.13", "TCGA.24.0982.01A.01R.1565.13",
                                "TCGA.24.1847.01A.01R.1566.13", "TCGA.24.2289.01A.01R.1568.13", "TCGA.31.1959.01A.01R.1568.13"))

result = df %>% separate(SAMPLE_ID,
                            into = c("Project", "TSS", "Participant", "Sample", "Vial",
                                     "Portion", "Analyte", "Plate", "Center"),
                            sep = "\\.")

错误输出:

Project TSS Participant Sample  Vial    Portion Analyte Plate   Center
<chr>   <chr>   <chr>   <chr>   <chr>   <chr>   <chr>   <chr>   <chr>
TCGA    13  1407    01A 01R 1565    13  NA  NA
TCGA    24  2254    01A 01R 1568    13  NA  NA
TCGA    24  0982    01A 01R 1565    13  NA  NA
TCGA    24  1847    01A 01R 1566    13  NA  NA
TCGA    24  2289    01A 01R 1568    13  NA  NA
TCGA    31  1959    01A 01R 1568    13  NA  NA

解决方案

方法1:分步拆分(清晰易读)

先按点拆分出基础字段,再拆分数字+字母的混合字段:

library(tidyr)
library(dplyr)

df <- data.frame(SAMPLE_ID = c("TCGA.13.1407.01A.01R.1565.13", 
                               "TCGA.24.2254.01A.01R.1568.13", 
                               "TCGA.24.0982.01A.01R.1565.13",
                               "TCGA.24.1847.01A.01R.1566.13", 
                               "TCGA.24.2289.01A.01R.1568.13", 
                               "TCGA.31.1959.01A.01R.1568.13"))

# 1. 按点拆分出7个临时字段
temp_df <- df %>%
  separate(SAMPLE_ID, 
           into = c("Project", "TSS", "Participant", "SampleVial", "PortionAnalyte", "Plate", "Center"),
           sep = "\\.")

# 2. 拆分SampleVial为Sample(数字)和Vial(字母)
temp_df <- temp_df %>%
  separate(SampleVial, into = c("Sample", "Vial"), sep = "(?<=\\d)(?=[A-Z])")

# 3. 拆分PortionAnalyte为Portion(数字)和Analyte(字母)
final_df <- temp_df %>%
  separate(PortionAnalyte, into = c("Portion", "Analyte"), sep = "(?<=\\d)(?=[A-Z])")

# 补充原SAMPLE_ID列并调整顺序
final_df <- final_df %>%
  mutate(SAMPLE_ID = df$SAMPLE_ID) %>%
  select(SAMPLE_ID, Project, TSS, Participant, Sample, Vial, Portion, Analyte, Plate, Center)

方法2:一次性拆分(简洁高效)

使用正则表达式同时匹配点分隔和数字-字母分隔,一次性完成拆分:

library(tidyr)
library(dplyr)

df <- data.frame(SAMPLE_ID = c("TCGA.13.1407.01A.01R.1565.13", 
                               "TCGA.24.2254.01A.01R.1568.13", 
                               "TCGA.24.0982.01A.01R.1565.13",
                               "TCGA.24.1847.01A.01R.1566.13", 
                               "TCGA.24.2289.01A.01R.1568.13", 
                               "TCGA.31.1959.01A.01R.1568.13"))

final_df <- df %>%
  separate(SAMPLE_ID,
           into = c("Project", "TSS", "Participant", "Sample", "Vial", "Portion", "Analyte", "Plate", "Center"),
           sep = "\\.|(?<=\\d)(?=[A-Z])")

结果验证

最终输出符合预期:

SAMPLE_ID Project TSS Participant Sample Vial Portion Analyte Plate Center
1 TCGA.13.1407.01A.01R.1565.13    TCGA  13        1407     01    A      01       R  1565     13
2 TCGA.24.2254.01A.01R.1568.13    TCGA  24        2254     01    A      01       R  1568     13
3 TCGA.24.0982.01A.01R.1565.13    TCGA  24        0982     01    A      01       R  1565     13
4 TCGA.24.1847.01A.01R.1566.13    TCGA  24        1847     01    A      01       R  1566     13
5 TCGA.24.2289.01A.01R.1568.13    TCGA  24        2289     01    A      01       R  1568     13
6 TCGA.31.1959.01A.01R.1568.13    TCGA  31        1959     01    A      01       R  1568     13

内容的提问来源于stack exchange,提问作者nicholaspooran

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.03 14:24:55