R语言数据框列拆分:从file列生成cell、insert、trt新列
解决R语言数据框拆分列并匹配参考数据的问题
方法一:利用参考数据框关联匹配
既然已经有包含完整映射关系的df2,可以先从file列提取cell字段,再通过cell与df2关联,直接获取对应的insert和trt:
library(tidyverse) # 原始数据 df <- tibble(file = c('U2_pN_Len', 'MM_pND_con', 'COS_CTL')) # 参考映射数据框 df2 <- tibble(cell = c('U2', 'MM', 'COS'), insert = c('pND', 'pN', 'pGFP'), trt = c('Len', 'con', 'CTL')) # 生成目标数据框 out <- df %>% # 提取file列中第一个下划线前的内容作为cell mutate(cell = str_extract(file, "^[^_]+")) %>% # 按cell字段关联参考数据框 left_join(df2, by = "cell") %>% # 调整列顺序与目标一致 select(file, cell, insert, trt) # 查看结果 out
运行后会得到期望的结果:
# A tibble: 3 × 4 file cell insert trt <chr> <chr> <chr> <chr> 1 U2_pN_Len U2 pND Len 2 MM_pND_con MM pN con 3 COS_CTL COS pGFP CTL
方法二:字符串拆分+手动映射(无需参考数据框)
如果不想依赖df2,可以先拆分file列,再通过case_when手动指定映射规则,同时处理拆分后列数不一致的情况:
library(tidyverse) df <- tibble(file = c('U2_pN_Len', 'MM_pND_con', 'COS_CTL')) out <- df %>% # 拆分file列,最多拆成3段,不足的补NA separate(file, into = c("cell", "temp", "trt"), sep = "_", fill = "right") %>% # 根据cell指定对应的insert值 mutate( insert = case_when( cell == "U2" ~ "pND", cell == "MM" ~ "pN", cell == "COS" ~ "pGFP", TRUE ~ NA_character_ ), # 处理trt列:当拆分后trt为NA时,用temp字段的值填充 trt = ifelse(is.na(trt), temp, trt) ) %>% select(file, cell, insert, trt) out
两种方法都能得到需要的最终数据框,推荐第一种方法,因为当映射关系需要修改时,只需调整df2即可,更易维护。
内容的提问来源于stack exchange,提问作者Abigail
相关产品推荐
相关产品推荐

