You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求Python/Bash脚本:按序列二级名称对FASTA序列去重

FASTA序列按二级名称去重的Python和Bash实现

刚好我之前处理过类似的FASTA去重需求,给你准备了两种简单易用的脚本方案,都能精准实现你要的逻辑:只要两条序列的二级名称(也就是FASTA ID行里第一个空格后面的内容)重复,就只保留第一次出现的那一条。


Python脚本实现

这个Python脚本可以处理多行的FASTA序列,而且支持从文件读取或者通过管道输入,灵活性很强。逻辑很直观:用一个集合记录已经见过的二级名称,遍历FASTA内容时,遇到新的二级名称就把对应的ID行和序列都输出,重复的直接跳过。

#!/usr/bin/env python3
import sys

def dedup_fasta(input_stream):
    seen_secondary = set()
    current_id = None
    current_seq = []
    
    for line in input_stream:
        line = line.strip()
        if not line:
            continue
        if line.startswith(">"):
            # 输出上一个收集完成的序列
            if current_id is not None:
                print(current_id)
                print(' '.join(current_seq))
            # 提取二级名称(第一个空格后的内容)
            parts = line.split(maxsplit=1)
            secondary_name = parts[1] if len(parts) >= 2 else parts[0]
            # 判断是否保留当前序列
            if secondary_name not in seen_secondary:
                seen_secondary.add(secondary_name)
                current_id = line
                current_seq = []
            else:
                # 跳过重复序列,重置状态
                current_id = None
                current_seq = []
        else:
            # 收集序列行(处理多行序列的情况)
            if current_id is not None:
                current_seq.append(line)
    # 输出最后一个序列
    if current_id is not None:
        print(current_id)
        print(' '.join(current_seq))

if __name__ == "__main__":
    if len(sys.argv) > 1:
        with open(sys.argv[1], 'r') as f:
            dedup_fasta(f)
    else:
        dedup_fasta(sys.stdin)

使用方法

  1. 把代码保存为dedup_fasta.py,赋予执行权限:chmod +x dedup_fasta.py
  2. 直接处理文件:./dedup_fasta.py input.fasta > output.fasta
  3. 或通过管道输入:cat input.fasta | ./dedup_fasta.py > output.fasta

Bash(Awk)脚本实现

如果你更习惯用命令行工具,用Awk写的脚本会更轻量,不需要依赖Python环境。同样基于记录已出现的二级名称实现去重,处理逻辑和Python版本一致。

#!/usr/bin/awk -f

BEGIN {
    delete seen
    current_seq = ""
    current_id = ""
}

/^>/ {
    # 输出上一个序列
    if (current_id != "") {
        print current_id
        print current_seq
    }
    # 提取二级名称
    if (index($0, " ") != 0) {
        secondary = substr($0, index($0, " ") + 1)
    } else {
        secondary = $0
    }
    # 判断是否保留当前序列
    if (!(secondary in seen)) {
        seen[secondary] = 1
        current_id = $0
        current_seq = ""
    } else {
        current_id = ""
        current_seq = ""
    }
    next
}

# 收集序列行
current_id != "" {
    current_seq = (current_seq == "") ? $0 : current_seq " " $0
}

END {
    # 输出最后一个序列
    if (current_id != "") {
        print current_id
        print current_seq
    }
}

使用方法

  1. 保存代码为dedup_fasta.awk,赋予执行权限:chmod +x dedup_fasta.awk
  2. 运行脚本:./dedup_fasta.awk input.fasta > output.fasta
  3. 或直接在命令行执行(无需保存文件):
awk '
BEGIN {
    delete seen
    current_seq = ""
    current_id = ""
}
/^>/ {
    if (current_id != "") {
        print current_id
        print current_seq
    }
    if (index($0, " ") != 0) {
        secondary = substr($0, index($0, " ") + 1)
    } else {
        secondary = $0
    }
    if (!(secondary in seen)) {
        seen[secondary] = 1
        current_id = $0
        current_seq = ""
    } else {
        current_id = ""
        current_seq = ""
    }
    next
}
current_id != "" {
    current_seq = (current_seq == "") ? $0 : current_seq " " $0
}
END {
    if (current_id != "") {
        print current_id
        print current_seq
    }
}' input.fasta > output.fasta

这两个脚本都能完美处理你提供的输入示例,输出结果和你预期的完全一致。如果遇到ID行没有空格的特殊情况,脚本会把整个ID作为标识来判断重复,避免遗漏或错误。

内容的提问来源于stack exchange,提问作者chippycentra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:45:14