You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

优化大体积PDBQT文件分割脚本:解决IO错误与效率问题

优化PDBQT大文件分割脚本:解决亿级分子IO报错与性能问题

你遇到的这个随机OSError(Errno 5)问题,本质是大流量IO操作下的系统资源波动或文件系统临时异常导致的,加上原脚本没有任何错误重试和异常处理机制,一旦出现单次IO失败就直接中断。另外原脚本的循环逻辑也有隐患,逐行遍历的效率在亿级数据下也会拖慢整体速度。我先梳理下原脚本的核心问题,再给出优化后的版本:

原脚本的核心问题

  • 没有IO异常捕获与重试,遇到文件系统临时错误直接崩溃
  • 循环逻辑存在漏洞:内层的break可能导致外层循环无法正确遍历所有分子
  • 频繁的文件打开/关闭操作,在大数量级下会加剧IO压力
  • 文件名处理没有做安全校验,遇到特殊字符可能导致重命名失败
  • 进度反馈不足,无法实时掌握处理状态

优化后的脚本

import os
import time
from pathlib import Path

def retry_io_operation(max_retries=3, delay=1):
    """IO操作重试装饰器,处理临时IO错误"""
    def decorator(func):
        def wrapper(*args, **kwargs):
            for attempt in range(max_retries):
                try:
                    return func(*args, **kwargs)
                except OSError as e:
                    if attempt == max_retries - 1:
                        raise
                    print(f"IO错误: {e},{delay}秒后重试...")
                    time.sleep(delay)
            return func(*args, **kwargs)
        return wrapper
    return decorator

@retry_io_operation(max_retries=5, delay=2)
def create_directory(dir_path):
    """安全创建目录,支持重试"""
    Path(dir_path).mkdir(parents=True, exist_ok=True)

@retry_io_operation(max_retries=5, delay=2)
def write_molecule_to_file(file_path, lines):
    """写入分子数据到文件,支持重试"""
    with open(file_path, 'w') as f:
        f.writelines(lines)

@retry_io_operation(max_retries=5, delay=2)
def rename_molecule_file(old_path, new_path):
    """重命名分子文件,支持重试"""
    if os.path.exists(new_path):
        # 处理重名情况,添加后缀避免覆盖
        base, ext = os.path.splitext(new_path)
        new_path = f"{base}_{int(time.time())}{ext}"
    os.rename(old_path, new_path)

def split_pdbqt(input_file, output_root, molecules_per_dir):
    input_path = Path(input_file)
    if not input_path.exists():
        raise FileNotFoundError(f"输入文件不存在: {input_file}")
    
    create_directory(output_root)
    
    molecule_count = 0
    dir_count = 1
    current_dir = os.path.join(output_root, str(dir_count))
    create_directory(current_dir)
    molecules_in_current_dir = 0
    
    with open(input_path, 'r') as infile:
        while True:
            # 定位到下一个MODEL块
            line = infile.readline()
            while line and not line.strip().startswith('MODEL '):
                line = infile.readline()
            if not line:
                break  # 文件读取完毕
            
            molecule_count += 1
            # 收集当前MODEL的所有行直到ENDMDL
            molecule_lines = [line]
            molecule_name = f"molecule_{molecule_count}"
            while True:
                line = infile.readline()
                if not line:
                    break
                molecule_lines.append(line)
                if line.strip() == 'ENDMDL':
                    break
                # 提取分子名称(如果有REMARK Name行)
                if line.strip().startswith('REMARK Name'):
                    parts = line.strip().split()
                    if len(parts) >=4:
                        molecule_name = parts[3]
                        # 清理非法字符,避免文件名错误
                        molecule_name = "".join(c for c in molecule_name if c not in '<>:"/\\|?*')
            
            # 写入临时文件,然后重命名
            temp_file = os.path.join(current_dir, f"temp_{molecule_count}.pdbqt")
            write_molecule_to_file(temp_file, molecule_lines)
            final_file = os.path.join(current_dir, f"{molecule_name}.pdbqt")
            rename_molecule_file(temp_file, final_file)
            
            molecules_in_current_dir += 1
            # 检查是否需要创建新目录
            if molecules_in_current_dir >= molecules_per_dir:
                print(f"[+] 完成目录 {current_dir},共 {molecules_in_current_dir} 个分子")
                dir_count += 1
                current_dir = os.path.join(output_root, str(dir_count))
                create_directory(current_dir)
                molecules_in_current_dir = 0
            
            # 每处理10000个分子打印进度,避免日志刷屏
            if molecule_count % 10000 == 0:
                print(f"[*] 已处理 {molecule_count} 个分子,当前目录: {current_dir}")
    
    print(f"\n----------\n[+] 处理完成!共分割 {molecule_count} 个分子,生成 {dir_count} 个目录")

# 调用示例:处理1亿个分子,每目录30万个
if __name__ == "__main__":
    split_pdbqt('L.pdbqt', 'Ligands', 300000)

关键优化说明

  • IO重试机制:通过装饰器给所有IO操作(创建目录、写文件、重命名)添加最多5次重试逻辑,遇到临时IO错误会自动重试,避免单次异常导致整个任务中断
  • 高效块读取:不再逐行遍历匹配MODEL,而是直接定位MODEL块后一次性读取整个分子的所有行,减少IO交互次数,提升处理速度
  • 文件名安全处理:清理分子名称中的非法字符,避免因特殊字符导致文件创建失败;同时处理重名情况,添加时间戳后缀防止文件覆盖
  • 进度可视化:每处理10000个分子打印进度,方便实时掌握任务状态
  • 目录预创建:提前创建当前目录,避免每次写入文件时重复检查目录存在性
  • 更健壮的循环逻辑:使用while True循环遍历文件,确保所有MODEL块都被处理,不会出现原脚本中内层break导致的遍历中断问题

这个脚本在处理亿级分子、大目录文件数的场景下,能大幅提升鲁棒性,同时因为优化了IO操作逻辑,处理效率也会比原脚本更高。

内容的提问来源于stack exchange,提问作者AC Research

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.07 22:57:55