求Python/Bash脚本:按序列二级名称对FASTA序列去重
FASTA序列按二级名称去重的Python和Bash实现
刚好我之前处理过类似的FASTA去重需求,给你准备了两种简单易用的脚本方案,都能精准实现你要的逻辑:只要两条序列的二级名称(也就是FASTA ID行里第一个空格后面的内容)重复,就只保留第一次出现的那一条。
Python脚本实现
这个Python脚本可以处理多行的FASTA序列,而且支持从文件读取或者通过管道输入,灵活性很强。逻辑很直观:用一个集合记录已经见过的二级名称,遍历FASTA内容时,遇到新的二级名称就把对应的ID行和序列都输出,重复的直接跳过。
#!/usr/bin/env python3 import sys def dedup_fasta(input_stream): seen_secondary = set() current_id = None current_seq = [] for line in input_stream: line = line.strip() if not line: continue if line.startswith(">"): # 输出上一个收集完成的序列 if current_id is not None: print(current_id) print(' '.join(current_seq)) # 提取二级名称(第一个空格后的内容) parts = line.split(maxsplit=1) secondary_name = parts[1] if len(parts) >= 2 else parts[0] # 判断是否保留当前序列 if secondary_name not in seen_secondary: seen_secondary.add(secondary_name) current_id = line current_seq = [] else: # 跳过重复序列,重置状态 current_id = None current_seq = [] else: # 收集序列行(处理多行序列的情况) if current_id is not None: current_seq.append(line) # 输出最后一个序列 if current_id is not None: print(current_id) print(' '.join(current_seq)) if __name__ == "__main__": if len(sys.argv) > 1: with open(sys.argv[1], 'r') as f: dedup_fasta(f) else: dedup_fasta(sys.stdin)
使用方法
- 把代码保存为
dedup_fasta.py,赋予执行权限:chmod +x dedup_fasta.py - 直接处理文件:
./dedup_fasta.py input.fasta > output.fasta - 或通过管道输入:
cat input.fasta | ./dedup_fasta.py > output.fasta
Bash(Awk)脚本实现
如果你更习惯用命令行工具,用Awk写的脚本会更轻量,不需要依赖Python环境。同样基于记录已出现的二级名称实现去重,处理逻辑和Python版本一致。
#!/usr/bin/awk -f BEGIN { delete seen current_seq = "" current_id = "" } /^>/ { # 输出上一个序列 if (current_id != "") { print current_id print current_seq } # 提取二级名称 if (index($0, " ") != 0) { secondary = substr($0, index($0, " ") + 1) } else { secondary = $0 } # 判断是否保留当前序列 if (!(secondary in seen)) { seen[secondary] = 1 current_id = $0 current_seq = "" } else { current_id = "" current_seq = "" } next } # 收集序列行 current_id != "" { current_seq = (current_seq == "") ? $0 : current_seq " " $0 } END { # 输出最后一个序列 if (current_id != "") { print current_id print current_seq } }
使用方法
- 保存代码为
dedup_fasta.awk,赋予执行权限:chmod +x dedup_fasta.awk - 运行脚本:
./dedup_fasta.awk input.fasta > output.fasta - 或直接在命令行执行(无需保存文件):
awk ' BEGIN { delete seen current_seq = "" current_id = "" } /^>/ { if (current_id != "") { print current_id print current_seq } if (index($0, " ") != 0) { secondary = substr($0, index($0, " ") + 1) } else { secondary = $0 } if (!(secondary in seen)) { seen[secondary] = 1 current_id = $0 current_seq = "" } else { current_id = "" current_seq = "" } next } current_id != "" { current_seq = (current_seq == "") ? $0 : current_seq " " $0 } END { if (current_id != "") { print current_id print current_seq } }' input.fasta > output.fasta
这两个脚本都能完美处理你提供的输入示例,输出结果和你预期的完全一致。如果遇到ID行没有空格的特殊情况,脚本会把整个ID作为标识来判断重复,避免遗漏或错误。
内容的提问来源于stack exchange,提问作者chippycentra
相关产品推荐
相关产品推荐

