You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python遍历Accession号的for循环无限运行,求排查原因

问题描述

我正在编写代码,从现有Excel文件中读取Accession号,通过Entrez工具执行BLAST分析。作为Python新手,我边学边写。当前代码持续运行却始终无法完成,也无错误返回。我曾用硬编码的单个Accession号和ID测试,代码可正常运行,因此问题出在循环设置上。

原代码
import pandas
from Bio import SeqIO
from Bio import Entrez
from Bio.Blast import NCBIWWW
from Bio.Blast import NCBIXML 

genesexcel=pandas.read_excel(r'C:\Users\bmn45\OneDrive - Drexel University\Monoterpenoid Indole Alkaloids\MIA_genes.xlsb', sheet_name='Data S16')  
#import the excel into variable

acc_num=genesexcel['Accession number'].values
#store values from first column of excel
print(acc_num)

for number in acc_num:

    handle = Entrez.esearch(db='protein', term=number, retmax="20")
    #searching the Genbank database using the accession number

    Entrez_result = Entrez.read(handle)
  
    ID_List = Entrez_result['IdList']
    #storing the ID number(s) from the Entrez search of the accession number

    handle2 = Entrez.efetch(db='protein', id=ID_List, rettype='gb')
    #uses efecth to get the full sequence record/file using the searched ID number
    #currently, the ID value is hardcoded, database is protein and the return file type (rettype) is set as a genbank file(gb)

    seq_record = SeqIO.read(handle2, "genbank")
    #storing the full sequence record as a genbank file
    seq_ID=(seq_record.name) 
    #storing the name of the sequence record (the full sequence ID)

    result_handle = NCBIWWW.qblast("tblastn", "nt", seq_ID, entrez_query="txid4056[ORGN]")
    #searching the nucleotide database ("nt") from protein sequence using tblastn, refined search related to Apocynacae (taxID4056)

    with open("my_blast_result.xml", "w") as out_handle:
        out_handle.write(result_handle.read())
    result_handle.close()
    #saving the blast results into a file

    result_handle = open("my_blast_result.xml")
    #loading saved BLAST results

    blast_records = NCBIXML.parse(result_handle)
    #parsing (analying) through the returned hits

    print("Alignments for sequence", next(blast_records).query)
    for alignment in blast_records:
        print("Accession number:", alignment.accession)
        print("Sequence:", alignment.title)
        print("Length:", alignment.length)
        print()
  
    print(result_handle.read())
问题定位与修复方案
  • 必须设置Entrez邮箱:NCBI要求使用Entrez工具时提供邮箱,否则会触发限流,导致请求无响应。添加Entrez.email = "your_email@example.com"(替换成你的真实邮箱)。
  • 避免结果文件被覆盖:循环中每次都写入同一个my_blast_result.xml,会覆盖之前的结果,还可能引发IO冲突。改为生成带Accession号的独立文件,比如blast_result_{number}.xml。
  • 正确处理Entrez连接:使用with语句自动关闭handle和handle2,避免连接泄漏。
  • 修复BLAST记录迭代问题:next(blast_records)会跳过第一条记录,后续循环只能处理剩余记录。建议先将所有记录转为列表,再统一处理。
  • 添加请求延迟:连续发送BLAST请求会被NCBI限流,每次循环后添加time.sleep(10)(可根据情况调整时长)。
  • 处理多序列返回情况:如果Entrez.efetch返回多条序列,SeqIO.read()会报错,改用SeqIO.parse()取第一条序列。
修改后的代码示例
import pandas
import time
from Bio import SeqIO
from Bio import Entrez
from Bio.Blast import NCBIWWW
from Bio.Blast import NCBIXML 

# 设置NCBI要求的邮箱
Entrez.email = "your_email@example.com"

genesexcel = pandas.read_excel(
    r'C:\Users\bmn45\OneDrive - Drexel University\Monoterpenoid Indole Alkaloids\MIA_genes.xlsb',
    sheet_name='Data S16'
)
acc_num = genesexcel['Accession number'].values
print(acc_num)

for number in acc_num:
    # 使用with自动关闭handle
    with Entrez.esearch(db='protein', term=number, retmax="20") as handle:
        Entrez_result = Entrez.read(handle)
        ID_List = Entrez_result['IdList']
    
    if not ID_List:
        print(f"未找到Accession号 {number} 对应的记录")
        continue
    
    # 获取序列记录,处理多序列情况
    with Entrez.efetch(db='protein', id=ID_List[0], rettype='gb') as handle2:
        # 取第一条序列
        seq_record = next(SeqIO.parse(handle2, "genbank"))
    seq_ID = seq_record.name
    print(f"正在处理序列: {seq_ID}")
    
    # 发送BLAST请求
    result_handle = NCBIWWW.qblast("tblastn", "nt", seq_ID, entrez_query="txid4056[ORGN]")
    
    # 保存为独立文件
    output_file = f"blast_result_{number}.xml"
    with open(output_file, "w") as out_handle:
        out_handle.write(result_handle.read())
    result_handle.close()
    
    # 读取并解析BLAST结果
    with open(output_file) as result_handle:
        blast_records = list(NCBIXML.parse(result_handle))
    
    if not blast_records:
        print(f"{number} 无BLAST结果")
        continue
    
    # 打印结果
    print(f"序列 {blast_records[0].query} 的比对结果:")
    for alignment in blast_records[0].alignments:
        print(f"Accession号: {alignment.accession}")
        print(f"序列标题: {alignment.title}")
        print(f"长度: {alignment.length}")
        print()
    
    # 添加延迟,避免限流
    time.sleep(10)

内容的提问来源于stack exchange,提问作者Brianna Nissley

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 14:43:13