You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过模式检测利用另一列替换字符列的值?附家系样本数据框示例

解决方案:基于模式匹配用另一列替换字符值

看起来你需要把fatherID和motherID中的短ID(比如AGG0093)替换成Sample_ID列对应的完整ID(比如UCSF_AGG0093_8005439857),结合你的数据结构,咱们用tidyverse工具链就能轻松实现,步骤如下:

1. 构建短ID到完整Sample_ID的映射表

首先从Sample_ID中提取出和fatherID/motherID匹配的短ID模式(比如你的数据里是AGG加数字的格式),然后建立两者的对应关系:

library(tidyverse)

# 提取Sample_ID中的短ID部分,创建映射关系
id_mapping <- pedigrees %>%
  mutate(short_id = str_extract(Sample_ID, "AGG\\d+")) %>%
  filter(short_id != "0") %>%  # 排除无父母的样本(ID为0的情况)
  select(short_id, Sample_ID) %>%
  deframe()  # 转换为named vector,方便后续替换

# 若你的tidyr版本没有deframe(),可以用这个替代写法:
# id_mapping <- pedigrees %>%
#   mutate(short_id = str_extract(Sample_ID, "AGG\\d+")) %>%
#   filter(short_id != "0") %>%
#   pull(Sample_ID, name = short_id)

2. 替换目标列的ID值

用上面的映射表,把fatherID和motherID里的短ID替换成完整Sample_ID,同时保留0(无父母标记)不变:

pedigrees_updated <- pedigrees %>%
  mutate(
    fatherID = ifelse(fatherID != "0", id_mapping[fatherID], fatherID),
    motherID = ifelse(motherID != "0", id_mapping[motherID], motherID)
  )

适配不同模式的小提示

如果你的短ID模式不是AGG\d+,只需要修改str_extract里的正则表达式即可:

  • 比如如果短ID是UCSF_AGG\d+前缀,就改成UCSF_AGG\\d+
  • 如果是任意数字字母组合的短ID,可尝试[A-Z0-9]+这类通用匹配规则

针对你给出的示例数据,替换后fatherID列的AGG0093会变成UCSF_AGG0093_8005439857,完全匹配对应的父母样本ID。


内容的提问来源于stack exchange,提问作者Carmen Sandoval

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:10:54