You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用DSSP分析蛋白质二级结构时遇FileNotFoundError求助

问题描述

我是Python新手,正在尝试用DSSP批量分析多个蛋白质的二级结构。脚本逻辑是针对每个UniProt编号从AlphaFold数据库下载对应PDB文件(路径为/Users/grb25/SS_HS/+uniprot+.pdb),之后调用Biopython的DSSP模块分析,但运行时一直触发FileNotFoundError: [WinError 2] The system cannot find the file specified报错。

报错信息

S-adenosylmethionine synthase isoform type-1
('/Users/grb25/SS_HS/Q00266.pdb', <http.client.HTTPMessage object at 0x000002A1ADE33D90>)
('/Users/grb25/SS_HS/Q00266.txt', <http.client.HTTPMessage object at 0x000002A1ADE33BE0>)
Q00266
The 'try' is finished
Traceback (most recent call last):
  File "<stdin>", line 31, in <module>
  File "C:\Users\grb25\AppData\Local\Programs\Python\Python310\lib\site-packages\Bio\PDB\DSSP.py", line 385, in __init__
    version_string = subprocess.check_output(
  File "C:\Users\grb25\AppData\Local\Programs\Python\Python310\lib\subprocess.py", line 421, in check_output
    return run(*popenargs, stdout=PIPE, timeout=timeout, check=True,
  File "C:\Users\grb25\AppData\Local\Programs\Python\Python310\lib\subprocess.py", line 503, in run
    with Popen(*popenargs, **kwargs) as process:
  File "C:\Users\grb25\AppData\Local\Programs\Python\Python310\lib\subprocess.py", line 971, in __init__
    self._execute_child(args, executable, preexec_fn, close_fds,
  File "C:\Users\grb25\AppData\Local\Programs\Python\Python310\lib\subprocess.py", line 1440, in _execute_child
    hp, ht, pid, tid = _winapi.CreateProcess(executable, args,
FileNotFoundError: [WinError 2] The system cannot find the file specified

核心代码片段

loc_pdb = "/Users/grb25/SS_HS/" + uniprot + ".pdb"
loc_error = "/Users/grb25/SS_HS/" + uniprot + ".txt"
struct = "https://alphafold.ebi.ac.uk/files/AF-{}-F1-model_v3.pdb".format(uniprot)
error = "https://alphafold.ebi.ac.uk/files/AF-{}-F1-model_v3.cif".format(uniprot)
urllib.request.urlretrieve(struct, loc_pdb)
urllib.request.urlretrieve(error, loc_error)
p = PDBParser()
structure = p.get_structure(uniprot, loc_pdb)
model = structure[0]
dssp = DSSP(model, "/Users/grb25/SS_HS/" + uniprot + ".pdb", dssp='mkhssp')

完整脚本片段

#For each unique protien accession number, go to the UniProt databse, and extract the protein's primary sequence
from Bio.PDB import PDBParser
from Bio.PDB.DSSP import DSSP
import urllib.request
index = 1
for uniprot in uniprots[1:]:    
    SS_list = []
    link = "https://www.uniprot.org/uniprot/" + uniprot +".fasta"
    f = session.get(link)
    contents = f.text.splitlines()
    try:
        #print(index)
        #print(contents)
        SS_list.append(uniprot)
        protein_name = ' '.join(contents[0].split("|")[-1].split("OS")[0].split()[1:])
        print (index)
        protein_seq = "".join("".join(contents[1:]))
        totalAA = len(protein_seq)
        SS_list.append(protein_name)
        SS_list.append(totalAA)
        print(protein_name)
        helix_count = 0
        sheet_count = 0
        turn_count = 0
        unstructured_count = 0
        loc_pdb = "/Users/grb25/SS_HS/" + uniprot + ".pdb"
        loc_error = "/Users/grb25/SS_HS/" + uniprot + ".txt"
        struct = "https://alphafold.ebi.ac.uk/files/AF-{}-F1-model_v3.pdb".format(uniprot)
        error = "https://alphafold.ebi.ac.uk/files/AF-{}-F1-model_v3.cif".format(uniprot)
        urllib.request.urlretrieve(struct, loc_pdb)
        urllib.request.urlretrieve(error, loc_error)
        p = PDBParser()
        structure = p.get_structure(uniprot, loc_pdb)
        model = structure[0]
        dssp = DSSP(model, "/Users/grb25/SS_HS/" + uniprot + ".pdb", dssp='mkhssp')
        SS_list.append(len(dssp.keys()))
        for i in range(len(dssp.keys())):
            a_key = list(dssp.keys())[i]
            if dssp[a_key][2]== "G" or dssp[a_key][2]== "H" or dssp[a_key][2]== "I":
                helix_count += 1
            elif dssp[a_key][2]== "E" or dssp[a_key][2]== "B":
                sheet_count += 1
            elif dssp[a_key][2]== "S" or dssp[a_key][2]== "T":
                turn_count += 1
            elif dssp[a_key][2]== "-":
                unstructured_count += 1
解决方法

这个报错的核心原因是你指定的dssp='mkhssp'可执行文件在Windows系统中找不到。Biopython的DSSP模块依赖外部的DSSP程序,可按以下方式解决:

  1. 安装官方DSSP程序并指定路径
    下载适用于Windows的DSSP程序,解压后记住其可执行文件的完整路径(比如C:\dssp\mkhssp.exe),修改DSSP初始化代码:

    dssp = DSSP(model, loc_pdb, dssp=r'C:\dssp\mkhssp.exe')
    

    用原始字符串(加r前缀)避免转义字符问题。

  2. 改用无外部依赖的内置DSSP实现
    如果你不想安装外部程序,可安装dssp Python包,然后修改代码去掉dssp参数:

    # 先执行pip install dssp安装依赖
    from Bio.PDB.DSSP import DSSP
    dssp = DSSP(model, loc_pdb)
    

    这种方式更适合新手,无需额外配置外部工具。

  3. 额外检查点

    • 确认下载的PDB文件确实存在于指定路径,且未损坏(可通过文本编辑器打开验证)。

内容的提问来源于stack exchange,提问作者Grace Bertles

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 15:34:57