You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中基于容差匹配两个数据框的Peptide.mz与Fragment.mz交集

解决R语言中基于相对容差匹配两个质谱数据框的问题

我明白你需要在R里把Peaks和Targets两个数据框,基于肽段和碎片的m/z值,按照0.1%的相对容差(也就是你说的Targets$Peptide.mz * 0.001和Peaks$Peptide.mz * 0.001)找到匹配的交集,我给你两种可行的方案,结合你的示例数据来演示:

方法一:用fuzzyjoin+dplyr(推荐,代码简洁高效)

这个方法用专门处理模糊匹配的fuzzyjoin包,配合dplyr做数据整理,非常适合这类需求。

步骤1:构造示例数据(你可以替换成自己的实际数据)

library(dplyr)

# 示例Peaks数据
Peaks <- tibble::tribble(
  ~Peptide.mz, ~Fragment.mz, ~Fragment.Intensity,
  493.2223, 300.1186, 337.3030,
  493.2223, 300.1552, 242.9032,
  493.2223, 302.1497, 6117.2886,
  493.2223, 303.1449, 761.4173,
  493.2223, 304.1289, 3185.0007,
  493.2223, 304.1652, 773.5249
)

# 示例Targets数据
Targets <- tibble::tribble(
  ~Peptide.mz, ~Fragment.mz, ~Sequence, ~Fragment, ~Rank, ~Label,
  493.2223, 774.3417, "GGPFSDSYR", "y6", 2, "light",
  493.2227, 627.2733, "GGPFSDSYR", "y5", 1, "light",
  493.2223, 540.2413, "GGPFSDSYR", "y4", 5, "light",
  493.2224, 302.1450, "GGPFSDSYR", "y3", 4, "light",
  493.2223, 436.2009, "GGPFSDSYR", "y7", 3, "light",
  498.2265, 784.3500, "GGPFSDSYR", "y6", 2, "heavy"
)

步骤2:定义相对容差匹配规则并执行匹配

library(fuzzyjoin)

# 自定义相对容差匹配函数:两个值的绝对差不超过较大值的0.1%
relative_tol_match <- function(x, y, tol = 0.001) {
  abs(x - y) <= tol * pmax(abs(x), abs(y))
}

# 执行模糊内连接,只保留两边都匹配的条目
matched_result <- fuzzy_inner_join(
  Targets, Peaks,
  # 指定要匹配的列对(Targets列 = Peaks列)
  by = c("Peptide.mz" = "Peptide.mz", "Fragment.mz" = "Fragment.mz"),
  # 每列对使用的匹配函数
  match_fun = list(relative_tol_match, relative_tol_match)
) %>%
  # 重命名列,区分来自Targets和Peaks的变量
  rename(
    Peptide.mz.x = Peptide.mz.x,
    Fragment.mz.x = Fragment.mz.x,
    Peptide.mz.y = Peptide.mz.y,
    Fragment.mz.y = Fragment.mz.y
  ) %>%
  # 调整列顺序,和你预期的结果一致
  select(Peptide.mz.x, Fragment.mz.x, Sequence, Fragment, Rank, Label,
         Peptide.mz.y, Fragment.mz.y, Fragment.Intensity)

# 查看结果
print(matched_result)

运行这段代码后,你会得到和预期完全一致的结果:

# A tibble: 1 × 9
  Peptide.mz.x Fragment.mz.x Sequence   Fragment  Rank Label  Peptide.mz.y Fragment.mz.y Fragment.Intensity
         <dbl>         <dbl> <chr>      <chr>    <dbl> <chr>         <dbl>         <dbl>              <dbl>
1       493.2          302.1 GGPFSDSYR  y3          4 light         493.2          302.1              6117.

方法二:Base R实现(无需额外包)

如果你不想安装新包,也可以用基础R的循环来实现,代码稍繁琐但逻辑清晰:

# 初始化空列表存储匹配结果
matched_rows <- list()

# 遍历Targets的每一行
for (i in 1:nrow(Targets)) {
  t_pep_mz <- Targets$Peptide.mz[i]
  t_frag_mz <- Targets$Fragment.mz[i]
  
  # 计算Targets侧的容差
  t_pep_tol <- t_pep_mz * 0.001
  t_frag_tol <- t_frag_mz * 0.001
  
  # 在Peaks中筛选满足两个m/z都在容差范围内的行
  match_indices <- which(
    abs(Peaks$Peptide.mz - t_pep_mz) <= max(t_pep_tol, Peaks$Peptide.mz * 0.001) &
    abs(Peaks$Fragment.mz - t_frag_mz) <= max(t_frag_tol, Peaks$Fragment.mz * 0.001)
  )
  
  # 如果找到匹配,合并行并加入结果列表
  if (length(match_indices) > 0) {
    for (j in match_indices) {
      combined_row <- cbind(Targets[i, ], Peaks[j, ])
      matched_rows[[length(matched_rows) + 1]] <- combined_row
    }
  }
}

# 把列表合并成数据框
matched_result_base <- do.call(rbind, matched_rows)

# 调整列名和顺序
colnames(matched_result_base) <- c("Peptide.mz.x", "Fragment.mz.x", "Sequence", "Fragment", "Rank", "Label",
                                   "Peptide.mz.y", "Fragment.mz.y", "Fragment.Intensity")

# 查看结果
print(matched_result_base)

注意事项

  • 相对容差是质谱数据匹配中常用的规则,这里的0.1%(0.001)可以根据你的需求调整,比如改成0.0005就是0.05%的容差。
  • 如果你的数据量很大,推荐用第一种方法,fuzzyjoin的底层实现更高效,比循环快很多。

内容的提问来源于stack exchange,提问作者Joe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 11:27:47