You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在R中匹配并合并.faa与.csv文件的相同元素

匹配CSV文件与FA蛋白序列文件并合并数据

我需要从CSV文件(data1_pt)的Protein.product列中扫描类似NP_000005.3的蛋白编号,在FA蛋白序列文件(data2_asfaa)中找到对应的匹配项,最终将匹配到的序列信息合并到原CSV表格中。


现有数据

1. CSV文件数据(data1_pt)

> head(data1_pt)
        X.Name    Accession  Start   Stop Strand GeneID    Locus Locus.tag Protein.product Length
1 chromosome 5 NC_000005.10  92232 182323      + 153478 PLEKHG4B         -     NP_443141.4   1627
2 chromosome 5 NC_000005.10 191539 195353      + 389257  LRRC14B         -  NP_001073947.1    514
3 chromosome 5 NC_000005.10 205297 216849      - 133957  CCDC127         -     NP_660308.1    260
4 chromosome 5 NC_000005.10 218356 263443      +   6389     SDHA         -  XP_011512374.1    680
5 chromosome 5 NC_000005.10 218356 263443      +   6389     SDHA         -  XP_011512375.1    599
6 chromosome 5 NC_000005.10 218356 263443      +   6389     SDHA         -  XP_047273423.1    632
                                              Protein.Name
1 pleckstrin homology domain-containing family G member 4B
2               leucine-rich repeat-containing protein 14B
3                coiled-coil domain-containing protein 127
4                                  succinate dehydrogenase
5                                  succinate dehydrogenase
6                                  succinate dehydrogenase

2. FA蛋白序列文件数据(data2_asfaa)

使用seqinr包读取:

library(seqinr)
data2_asfaa <- read.fasta("GCF_000001405.40_GRCh38.p14_protein.faa")

序列摘要:

> summary(data2_asfaa)
               Length Class       Mode     
NP_000005.3     1474  SeqFastadna character
NP_000006.2      290  SeqFastadna character
NP_000007.1      421  SeqFastadna character
NP_000008.1      412  SeqFastadna character
NP_000009.1      655  SeqFastadna character
NP_000010.1      427  SeqFastadna character
NP_000011.2      503  SeqFastadna character
NP_000012.1      467  SeqFastadna character
NP_000013.2      363  SeqFastadna character
NP_000014.1      387  SeqFastadna character
NP_000015.2      413  SeqFastadna character
NP_000016.1      408  SeqFastadna character
NP_000017.1      484  SeqFastadna character
NP_000018.2      346  SeqFastadna character
NP_000019.2     1532  SeqFastadna character

单个序列详情:

> head(data2_asfaa)
$NP_000005.3
   [1] "m" "g" "k" "n" "k" "l" "l" "h" "p" "s" "l" "v" "l" "l" "l" "l" "v" "l" "l" "p" "t" "d" "a" "s" "v" "s" "g" "k"
  [29] "p" "q" "y" "m" "v" "l" "v" "p" "s" "l" "l" "h" "t" "e" "t" "t" "e" "k" "g" "c" "v" "l" "l" "s" "y" "l" "n" "e"
  [57] "t" "v" "t" "v" "s" "a" "s" "l" "e" "s" "v" "r" "g" "n" "r" "s" "l" "f" "t" "d" "l" "e" "a" "e" "n" "d" "v" "l"
  [85] "h" "c" "v" "a" "f" "a" "v" "p" "k" "s" "s" "s" "n" "e" "e" "v" "m" "f" "l" "t" "v" "q" "v" "k" "g" "p" "t" "q"
 [113] "e" "f" "k" "k" "r" "t" "t" "v" "m" "v" "k" "n" "e" "d" "s" "l" "v" "f" "v" "q" "t" "d" "k" "s" "i" "y" "k" "p"
 [141] "g" "q" "t" "v" "k" "f" "r" "v" "v" "s" "m" "d" "e" "n" "f" "h" "p" "l" "n" "e" "l" "i" "p" "l" "v" "y" "i" "q"
 [169] "d" "p" "k" "g" "n" "r" "i" "a" "q" "w" "q" "s" "f" "q" "l" "e" "g" "g" "l" "k" "q" "f" "s" "f" "p" "l" "s" "s"
   [ reached getOption("max.print") -- omitted 474 entries ]
    attr(,"name")
    [1] "NP_000005.3"
    attr(,"Annot")
    [1] ">NP_000005.3 alpha-2-macroglobulin isoform a precursor [Homo sapiens]"
    attr(,"class")
    [1] "SeqFastadna"

我尝试过但无效的代码

thing = ""
for (i in length(data2_asfaa)){
  thing <- data1_pt == data2_asfaa& #can't write a column name because its not tabular
positions <- c(which(sapply(pproducts, list(data2_asfaa), pproducts %in% data2_asfaa)))

正确解决方案

步骤1:整理FA文件数据为数据框

将FA文件中的蛋白编号、完整序列、注释信息提取出来,转换成匹配CSV格式的数据框:

# 提取蛋白编号(即列表元素的名称)
protein_ids <- names(data2_asfaa)
# 将氨基酸字符向量拼接成完整序列字符串
protein_seqs <- sapply(data2_asfaa, function(x) paste(x, collapse = ""))
# 提取每个序列的注释信息
protein_annotations <- sapply(data2_asfaa, function(x) attr(x, "Annot"))

# 组合成标准数据框
faa_df <- data.frame(
  Protein.product = protein_ids,
  Protein_Sequence = protein_seqs,
  Protein_Annotation = protein_annotations,
  stringsAsFactors = FALSE
)

步骤2:合并两个数据集

以Protein.product为匹配键,将原CSV数据与FA数据框合并:

# 保留原CSV的所有行,匹配不到的位置填充NA
merged_data <- merge(data1_pt, faa_df, by = "Protein.product", all.x = TRUE)

# 若只需要保留两边都匹配的记录,去掉all.x参数即可(默认内连接)
# merged_data <- merge(data1_pt, faa_df, by = "Protein.product")

步骤3:查看合并结果

# 输出前几行合并后的数据
head(merged_data)

内容的提问来源于stack exchange,提问作者Ayaz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 09:50:42