You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何拼接multiFASTA文件内序列并合并输出为保留头信息的单个FASTA文件

解决方案

方案1:Linux命令行(awk实现,无额外依赖,适配zsh环境)

原有循环出错核心原因有两点:1. 循环内使用cat *.fasta会读取目录下全部fasta文件而非当前遍历的单个文件;2. 每次使用>覆盖输出文件,导致仅保留最后一次处理结果。
以下代码直接实现需求:

# 先清空目标结果文件
> GENE_L.fasta
# 遍历当前目录下所有后缀为.fasta的文件
for f in *.fasta; do
    awk '
    # 匹配当前文件的第一个序列头,直接输出
    FNR==1 {print $0}
    # 匹配非序列头的序列行,追加到拼接变量
    !/^>/ {concat_seq = concat_seq $0}
    # 单个文件处理完成后输出拼接好的序列,清空变量供下一个文件使用
    END {print concat_seq; concat_seq=""}
    ' "$f" >> GENE_L.fasta
done

方案2:Python实现(依赖Biopython,格式兼容性好,易扩展)

如果需要更高的格式容错性、后续方便扩展处理逻辑,可以使用Biopython实现:
首先安装依赖:
pip install biopython
运行代码:

from Bio import SeqIO
from pathlib import Path

output_file = "GENE_L.fasta"
processed_seqs = []

# 遍历所有fasta文件
for fasta_path in Path(".").glob("*.fasta"):
    records = list(SeqIO.parse(fasta_path, "fasta"))
    if not records:
        continue
    # 保留第一个序列的ID和注释,拼接所有序列
    merged_record = records[0]
    merged_record.seq = sum([rec.seq for rec in records], merged_record.seq.__class__(""))
    processed_seqs.append(merged_record)

# 写入最终合并文件
SeqIO.write(processed_seqs, output_file, "fasta")

内容的提问来源于stack exchange,提问作者Fitzy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.01 09:36:02