Python遍历Accession号的for循环无限运行,求排查原因
问题描述
我正在编写代码,从现有Excel文件中读取Accession号,通过Entrez工具执行BLAST分析。作为Python新手,我边学边写。当前代码持续运行却始终无法完成,也无错误返回。我曾用硬编码的单个Accession号和ID测试,代码可正常运行,因此问题出在循环设置上。
原代码
import pandas from Bio import SeqIO from Bio import Entrez from Bio.Blast import NCBIWWW from Bio.Blast import NCBIXML genesexcel=pandas.read_excel(r'C:\Users\bmn45\OneDrive - Drexel University\Monoterpenoid Indole Alkaloids\MIA_genes.xlsb', sheet_name='Data S16') #import the excel into variable acc_num=genesexcel['Accession number'].values #store values from first column of excel print(acc_num) for number in acc_num: handle = Entrez.esearch(db='protein', term=number, retmax="20") #searching the Genbank database using the accession number Entrez_result = Entrez.read(handle) ID_List = Entrez_result['IdList'] #storing the ID number(s) from the Entrez search of the accession number handle2 = Entrez.efetch(db='protein', id=ID_List, rettype='gb') #uses efecth to get the full sequence record/file using the searched ID number #currently, the ID value is hardcoded, database is protein and the return file type (rettype) is set as a genbank file(gb) seq_record = SeqIO.read(handle2, "genbank") #storing the full sequence record as a genbank file seq_ID=(seq_record.name) #storing the name of the sequence record (the full sequence ID) result_handle = NCBIWWW.qblast("tblastn", "nt", seq_ID, entrez_query="txid4056[ORGN]") #searching the nucleotide database ("nt") from protein sequence using tblastn, refined search related to Apocynacae (taxID4056) with open("my_blast_result.xml", "w") as out_handle: out_handle.write(result_handle.read()) result_handle.close() #saving the blast results into a file result_handle = open("my_blast_result.xml") #loading saved BLAST results blast_records = NCBIXML.parse(result_handle) #parsing (analying) through the returned hits print("Alignments for sequence", next(blast_records).query) for alignment in blast_records: print("Accession number:", alignment.accession) print("Sequence:", alignment.title) print("Length:", alignment.length) print() print(result_handle.read())
问题定位与修复方案
- 必须设置Entrez邮箱:NCBI要求使用Entrez工具时提供邮箱,否则会触发限流,导致请求无响应。添加
Entrez.email = "your_email@example.com"(替换成你的真实邮箱)。 - 避免结果文件被覆盖:循环中每次都写入同一个
my_blast_result.xml,会覆盖之前的结果,还可能引发IO冲突。改为生成带Accession号的独立文件,比如blast_result_{number}.xml。 - 正确处理Entrez连接:使用
with语句自动关闭handle和handle2,避免连接泄漏。 - 修复BLAST记录迭代问题:
next(blast_records)会跳过第一条记录,后续循环只能处理剩余记录。建议先将所有记录转为列表,再统一处理。 - 添加请求延迟:连续发送BLAST请求会被NCBI限流,每次循环后添加
time.sleep(10)(可根据情况调整时长)。 - 处理多序列返回情况:如果
Entrez.efetch返回多条序列,SeqIO.read()会报错,改用SeqIO.parse()取第一条序列。
修改后的代码示例
import pandas import time from Bio import SeqIO from Bio import Entrez from Bio.Blast import NCBIWWW from Bio.Blast import NCBIXML # 设置NCBI要求的邮箱 Entrez.email = "your_email@example.com" genesexcel = pandas.read_excel( r'C:\Users\bmn45\OneDrive - Drexel University\Monoterpenoid Indole Alkaloids\MIA_genes.xlsb', sheet_name='Data S16' ) acc_num = genesexcel['Accession number'].values print(acc_num) for number in acc_num: # 使用with自动关闭handle with Entrez.esearch(db='protein', term=number, retmax="20") as handle: Entrez_result = Entrez.read(handle) ID_List = Entrez_result['IdList'] if not ID_List: print(f"未找到Accession号 {number} 对应的记录") continue # 获取序列记录,处理多序列情况 with Entrez.efetch(db='protein', id=ID_List[0], rettype='gb') as handle2: # 取第一条序列 seq_record = next(SeqIO.parse(handle2, "genbank")) seq_ID = seq_record.name print(f"正在处理序列: {seq_ID}") # 发送BLAST请求 result_handle = NCBIWWW.qblast("tblastn", "nt", seq_ID, entrez_query="txid4056[ORGN]") # 保存为独立文件 output_file = f"blast_result_{number}.xml" with open(output_file, "w") as out_handle: out_handle.write(result_handle.read()) result_handle.close() # 读取并解析BLAST结果 with open(output_file) as result_handle: blast_records = list(NCBIXML.parse(result_handle)) if not blast_records: print(f"{number} 无BLAST结果") continue # 打印结果 print(f"序列 {blast_records[0].query} 的比对结果:") for alignment in blast_records[0].alignments: print(f"Accession号: {alignment.accession}") print(f"序列标题: {alignment.title}") print(f"长度: {alignment.length}") print() # 添加延迟,避免限流 time.sleep(10)
内容的提问来源于stack exchange,提问作者Brianna Nissley
相关产品推荐
相关产品推荐

