如何通过模式检测利用另一列替换字符列的值?附家系样本数据框示例
解决方案:基于模式匹配用另一列替换字符值
看起来你需要把fatherID和motherID中的短ID(比如AGG0093)替换成Sample_ID列对应的完整ID(比如UCSF_AGG0093_8005439857),结合你的数据结构,咱们用tidyverse工具链就能轻松实现,步骤如下:
1. 构建短ID到完整Sample_ID的映射表
首先从Sample_ID中提取出和fatherID/motherID匹配的短ID模式(比如你的数据里是AGG加数字的格式),然后建立两者的对应关系:
library(tidyverse) # 提取Sample_ID中的短ID部分,创建映射关系 id_mapping <- pedigrees %>% mutate(short_id = str_extract(Sample_ID, "AGG\\d+")) %>% filter(short_id != "0") %>% # 排除无父母的样本(ID为0的情况) select(short_id, Sample_ID) %>% deframe() # 转换为named vector,方便后续替换 # 若你的tidyr版本没有deframe(),可以用这个替代写法: # id_mapping <- pedigrees %>% # mutate(short_id = str_extract(Sample_ID, "AGG\\d+")) %>% # filter(short_id != "0") %>% # pull(Sample_ID, name = short_id)
2. 替换目标列的ID值
用上面的映射表,把fatherID和motherID里的短ID替换成完整Sample_ID,同时保留0(无父母标记)不变:
pedigrees_updated <- pedigrees %>% mutate( fatherID = ifelse(fatherID != "0", id_mapping[fatherID], fatherID), motherID = ifelse(motherID != "0", id_mapping[motherID], motherID) )
适配不同模式的小提示
如果你的短ID模式不是AGG\d+,只需要修改str_extract里的正则表达式即可:
- 比如如果短ID是
UCSF_AGG\d+前缀,就改成UCSF_AGG\\d+ - 如果是任意数字字母组合的短ID,可尝试
[A-Z0-9]+这类通用匹配规则
针对你给出的示例数据,替换后fatherID列的AGG0093会变成UCSF_AGG0093_8005439857,完全匹配对应的父母样本ID。
内容的提问来源于stack exchange,提问作者Carmen Sandoval
相关产品推荐
相关产品推荐

