You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从PDB文件中提取chain-ID并生成指定格式的纯文本文件?

操作步骤

前置依赖

首先安装Biopython用于PDB文件解析,执行命令:
pip install biopython

核心逻辑

  • 遍历本地所有符合pdbXXXX.ent.gz命名格式的文件,从文件名提取4位PDB ID
  • 直接读取压缩格式的PDB文件,无需手动解压,提取文件头存储的分辨率信息,以及结构中的所有链ID
  • 将每一条PDB ID、链ID、分辨率的对应记录按要求格式写入输出文本

可直接运行的代码示例

import os
from Bio.PDB import PDBParser
from Bio.PDB.PDBExceptions import PDBConstructionWarning
import warnings

# 忽略PDB解析过程中的非致命结构警告
warnings.filterwarnings("ignore", category=PDBConstructionWarning)

# 请修改以下两个路径为你自己的本地路径
PDB_FILE_DIR = "/your/local/pdb/storage/path"
OUTPUT_TXT_PATH = "./pdb_chain_res_list.txt"

parser = PDBParser(QUIET=True)

with open(OUTPUT_TXT_PATH, "w", encoding="utf-8") as out_f:
    # 如果你不需要表头可以注释掉下一行
    out_f.write("pdb_id\tchain_id\tresolution\n")
    for fname in os.listdir(PDB_FILE_DIR):
        # 过滤不符合命名格式的文件
        if not fname.startswith("pdb") or not fname.endswith(".ent.gz"):
            continue
        pdb_id = fname[3:7].lower()
        full_path = os.path.join(PDB_FILE_DIR, fname)
        try:
            struct = parser.get_structure(pdb_id, full_path)
            # 提取分辨率,无分辨率则跳过
            res = struct.header.get("resolution")
            if not res:
                continue
            # 遍历所有链
            for model in struct:
                for chain in model:
                    chain_id = chain.get_id()
                    out_f.write(f"{pdb_id}\t{chain_id}\t{res:.2f}\n")
        except:
            # 解析失败的文件直接跳过,可自行添加日志输出排查问题
            continue

常见调整需求

  • 如果你的PDB文件存放在多层子目录中,将os.listdir遍历替换为os.walk递归遍历即可
  • 如果需要保留无分辨率的记录,将if not res: continue改为res = "N/A"即可
  • 如果要求用空格而非制表符分隔,将输出语句中的\t替换为对应数量的空格即可

内容的提问来源于stack exchange,提问作者user366312

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.27 22:45:03