读取FASTA序列时遗漏最后一条序列的问题求助
为什么我的FASTA读取代码总是跳过最后一条序列?
嘿,我一眼就瞅出问题啦!你的代码逻辑是只有遇到下一个>开头的标题行时,才会把当前攒的序列存到test里,但最后一条序列读完之后,根本没有后续的标题行触发这个保存操作,所以它就一直留在fasta列表里被彻底遗漏了。
先看看你原来的代码:
import sys fasta = [] test = [] with open("tests.fasta") as file_one: for line in file_one: line = line.strip() if not line: continue if line.startswith(">"): active_sequence_name = line[1:] if active_sequence_name not in fasta: test.append(''.join(fasta)) fasta = [] continue sequence = line fasta.append(sequence) print(...)
这里还有个小bug:if active_sequence_name not in fasta完全逻辑错误——fasta是用来存序列片段的列表,不是存标题的,这行根本起不到你想要的作用,反而可能导致莫名其妙的问题。
给你修正后的代码,解决这两个问题:
import sys fasta = [] test = [] with open("tests.fasta") as file_one: for line in file_one: line = line.strip() if not line: continue if line.startswith(">"): # 只有当当前有正在处理的序列时,才保存到test if fasta: test.append(''.join(fasta)) fasta = [] # 这里可以记录标题,如果需要的话 active_sequence_name = line[1:] continue sequence = line fasta.append(sequence) # 关键:循环结束后,别忘了处理最后一条还没保存的序列 if fasta: test.append(''.join(fasta)) # 现在打印test就能看到所有序列了 print(test)
简单说下修正点:
- 删掉了逻辑错误的
active_sequence_name not in fasta判断,改成检查fasta是否非空——只有当我们已经攒了一些序列片段时,才需要保存并清空列表。 - 在循环完全结束后,专门检查一次
fasta列表,如果里面还有内容(也就是最后一条序列),就把它合并后加入test。这一步就是解决你遗漏最后一条序列的核心!
内容的提问来源于stack exchange,提问作者Nitin
相关产品推荐
相关产品推荐

