在R中匹配并合并.faa与.csv文件的相同元素
匹配CSV文件与FA蛋白序列文件并合并数据
我需要从CSV文件(data1_pt)的Protein.product列中扫描类似NP_000005.3的蛋白编号,在FA蛋白序列文件(data2_asfaa)中找到对应的匹配项,最终将匹配到的序列信息合并到原CSV表格中。
现有数据
1. CSV文件数据(data1_pt)
> head(data1_pt) X.Name Accession Start Stop Strand GeneID Locus Locus.tag Protein.product Length 1 chromosome 5 NC_000005.10 92232 182323 + 153478 PLEKHG4B - NP_443141.4 1627 2 chromosome 5 NC_000005.10 191539 195353 + 389257 LRRC14B - NP_001073947.1 514 3 chromosome 5 NC_000005.10 205297 216849 - 133957 CCDC127 - NP_660308.1 260 4 chromosome 5 NC_000005.10 218356 263443 + 6389 SDHA - XP_011512374.1 680 5 chromosome 5 NC_000005.10 218356 263443 + 6389 SDHA - XP_011512375.1 599 6 chromosome 5 NC_000005.10 218356 263443 + 6389 SDHA - XP_047273423.1 632 Protein.Name 1 pleckstrin homology domain-containing family G member 4B 2 leucine-rich repeat-containing protein 14B 3 coiled-coil domain-containing protein 127 4 succinate dehydrogenase 5 succinate dehydrogenase 6 succinate dehydrogenase
2. FA蛋白序列文件数据(data2_asfaa)
使用seqinr包读取:
library(seqinr) data2_asfaa <- read.fasta("GCF_000001405.40_GRCh38.p14_protein.faa")
序列摘要:
> summary(data2_asfaa) Length Class Mode NP_000005.3 1474 SeqFastadna character NP_000006.2 290 SeqFastadna character NP_000007.1 421 SeqFastadna character NP_000008.1 412 SeqFastadna character NP_000009.1 655 SeqFastadna character NP_000010.1 427 SeqFastadna character NP_000011.2 503 SeqFastadna character NP_000012.1 467 SeqFastadna character NP_000013.2 363 SeqFastadna character NP_000014.1 387 SeqFastadna character NP_000015.2 413 SeqFastadna character NP_000016.1 408 SeqFastadna character NP_000017.1 484 SeqFastadna character NP_000018.2 346 SeqFastadna character NP_000019.2 1532 SeqFastadna character
单个序列详情:
> head(data2_asfaa) $NP_000005.3 [1] "m" "g" "k" "n" "k" "l" "l" "h" "p" "s" "l" "v" "l" "l" "l" "l" "v" "l" "l" "p" "t" "d" "a" "s" "v" "s" "g" "k" [29] "p" "q" "y" "m" "v" "l" "v" "p" "s" "l" "l" "h" "t" "e" "t" "t" "e" "k" "g" "c" "v" "l" "l" "s" "y" "l" "n" "e" [57] "t" "v" "t" "v" "s" "a" "s" "l" "e" "s" "v" "r" "g" "n" "r" "s" "l" "f" "t" "d" "l" "e" "a" "e" "n" "d" "v" "l" [85] "h" "c" "v" "a" "f" "a" "v" "p" "k" "s" "s" "s" "n" "e" "e" "v" "m" "f" "l" "t" "v" "q" "v" "k" "g" "p" "t" "q" [113] "e" "f" "k" "k" "r" "t" "t" "v" "m" "v" "k" "n" "e" "d" "s" "l" "v" "f" "v" "q" "t" "d" "k" "s" "i" "y" "k" "p" [141] "g" "q" "t" "v" "k" "f" "r" "v" "v" "s" "m" "d" "e" "n" "f" "h" "p" "l" "n" "e" "l" "i" "p" "l" "v" "y" "i" "q" [169] "d" "p" "k" "g" "n" "r" "i" "a" "q" "w" "q" "s" "f" "q" "l" "e" "g" "g" "l" "k" "q" "f" "s" "f" "p" "l" "s" "s" [ reached getOption("max.print") -- omitted 474 entries ] attr(,"name") [1] "NP_000005.3" attr(,"Annot") [1] ">NP_000005.3 alpha-2-macroglobulin isoform a precursor [Homo sapiens]" attr(,"class") [1] "SeqFastadna"
我尝试过但无效的代码
thing = "" for (i in length(data2_asfaa)){ thing <- data1_pt == data2_asfaa& #can't write a column name because its not tabular
positions <- c(which(sapply(pproducts, list(data2_asfaa), pproducts %in% data2_asfaa)))
正确解决方案
步骤1:整理FA文件数据为数据框
将FA文件中的蛋白编号、完整序列、注释信息提取出来,转换成匹配CSV格式的数据框:
# 提取蛋白编号(即列表元素的名称) protein_ids <- names(data2_asfaa) # 将氨基酸字符向量拼接成完整序列字符串 protein_seqs <- sapply(data2_asfaa, function(x) paste(x, collapse = "")) # 提取每个序列的注释信息 protein_annotations <- sapply(data2_asfaa, function(x) attr(x, "Annot")) # 组合成标准数据框 faa_df <- data.frame( Protein.product = protein_ids, Protein_Sequence = protein_seqs, Protein_Annotation = protein_annotations, stringsAsFactors = FALSE )
步骤2:合并两个数据集
以Protein.product为匹配键,将原CSV数据与FA数据框合并:
# 保留原CSV的所有行,匹配不到的位置填充NA merged_data <- merge(data1_pt, faa_df, by = "Protein.product", all.x = TRUE) # 若只需要保留两边都匹配的记录,去掉all.x参数即可(默认内连接) # merged_data <- merge(data1_pt, faa_df, by = "Protein.product")
步骤3:查看合并结果
# 输出前几行合并后的数据 head(merged_data)
内容的提问来源于stack exchange,提问作者Ayaz
相关产品推荐
相关产品推荐

