优化大体积PDBQT文件分割脚本:解决IO错误与效率问题
优化PDBQT大文件分割脚本:解决亿级分子IO报错与性能问题
你遇到的这个随机OSError(Errno 5)问题,本质是大流量IO操作下的系统资源波动或文件系统临时异常导致的,加上原脚本没有任何错误重试和异常处理机制,一旦出现单次IO失败就直接中断。另外原脚本的循环逻辑也有隐患,逐行遍历的效率在亿级数据下也会拖慢整体速度。我先梳理下原脚本的核心问题,再给出优化后的版本:
原脚本的核心问题
- 没有IO异常捕获与重试,遇到文件系统临时错误直接崩溃
- 循环逻辑存在漏洞:内层的
break可能导致外层循环无法正确遍历所有分子 - 频繁的文件打开/关闭操作,在大数量级下会加剧IO压力
- 文件名处理没有做安全校验,遇到特殊字符可能导致重命名失败
- 进度反馈不足,无法实时掌握处理状态
优化后的脚本
import os import time from pathlib import Path def retry_io_operation(max_retries=3, delay=1): """IO操作重试装饰器,处理临时IO错误""" def decorator(func): def wrapper(*args, **kwargs): for attempt in range(max_retries): try: return func(*args, **kwargs) except OSError as e: if attempt == max_retries - 1: raise print(f"IO错误: {e},{delay}秒后重试...") time.sleep(delay) return func(*args, **kwargs) return wrapper return decorator @retry_io_operation(max_retries=5, delay=2) def create_directory(dir_path): """安全创建目录,支持重试""" Path(dir_path).mkdir(parents=True, exist_ok=True) @retry_io_operation(max_retries=5, delay=2) def write_molecule_to_file(file_path, lines): """写入分子数据到文件,支持重试""" with open(file_path, 'w') as f: f.writelines(lines) @retry_io_operation(max_retries=5, delay=2) def rename_molecule_file(old_path, new_path): """重命名分子文件,支持重试""" if os.path.exists(new_path): # 处理重名情况,添加后缀避免覆盖 base, ext = os.path.splitext(new_path) new_path = f"{base}_{int(time.time())}{ext}" os.rename(old_path, new_path) def split_pdbqt(input_file, output_root, molecules_per_dir): input_path = Path(input_file) if not input_path.exists(): raise FileNotFoundError(f"输入文件不存在: {input_file}") create_directory(output_root) molecule_count = 0 dir_count = 1 current_dir = os.path.join(output_root, str(dir_count)) create_directory(current_dir) molecules_in_current_dir = 0 with open(input_path, 'r') as infile: while True: # 定位到下一个MODEL块 line = infile.readline() while line and not line.strip().startswith('MODEL '): line = infile.readline() if not line: break # 文件读取完毕 molecule_count += 1 # 收集当前MODEL的所有行直到ENDMDL molecule_lines = [line] molecule_name = f"molecule_{molecule_count}" while True: line = infile.readline() if not line: break molecule_lines.append(line) if line.strip() == 'ENDMDL': break # 提取分子名称(如果有REMARK Name行) if line.strip().startswith('REMARK Name'): parts = line.strip().split() if len(parts) >=4: molecule_name = parts[3] # 清理非法字符,避免文件名错误 molecule_name = "".join(c for c in molecule_name if c not in '<>:"/\\|?*') # 写入临时文件,然后重命名 temp_file = os.path.join(current_dir, f"temp_{molecule_count}.pdbqt") write_molecule_to_file(temp_file, molecule_lines) final_file = os.path.join(current_dir, f"{molecule_name}.pdbqt") rename_molecule_file(temp_file, final_file) molecules_in_current_dir += 1 # 检查是否需要创建新目录 if molecules_in_current_dir >= molecules_per_dir: print(f"[+] 完成目录 {current_dir},共 {molecules_in_current_dir} 个分子") dir_count += 1 current_dir = os.path.join(output_root, str(dir_count)) create_directory(current_dir) molecules_in_current_dir = 0 # 每处理10000个分子打印进度,避免日志刷屏 if molecule_count % 10000 == 0: print(f"[*] 已处理 {molecule_count} 个分子,当前目录: {current_dir}") print(f"\n----------\n[+] 处理完成!共分割 {molecule_count} 个分子,生成 {dir_count} 个目录") # 调用示例:处理1亿个分子,每目录30万个 if __name__ == "__main__": split_pdbqt('L.pdbqt', 'Ligands', 300000)
关键优化说明
- IO重试机制:通过装饰器给所有IO操作(创建目录、写文件、重命名)添加最多5次重试逻辑,遇到临时IO错误会自动重试,避免单次异常导致整个任务中断
- 高效块读取:不再逐行遍历匹配MODEL,而是直接定位MODEL块后一次性读取整个分子的所有行,减少IO交互次数,提升处理速度
- 文件名安全处理:清理分子名称中的非法字符,避免因特殊字符导致文件创建失败;同时处理重名情况,添加时间戳后缀防止文件覆盖
- 进度可视化:每处理10000个分子打印进度,方便实时掌握任务状态
- 目录预创建:提前创建当前目录,避免每次写入文件时重复检查目录存在性
- 更健壮的循环逻辑:使用
while True循环遍历文件,确保所有MODEL块都被处理,不会出现原脚本中内层break导致的遍历中断问题
这个脚本在处理亿级分子、大目录文件数的场景下,能大幅提升鲁棒性,同时因为优化了IO操作逻辑,处理效率也会比原脚本更高。
内容的提问来源于stack exchange,提问作者AC Research
相关产品推荐
相关产品推荐

