批量处理pdbqt文件:基于TORSDOF出现情况删除冗余行
问题
我有7000个pdbqt格式文件(从sade1.pdbqt到sade7200.pdbqt),部分文件里关键词TORSDOF会出现第二次甚至多次。需要对这类文件做如下处理:
- 如果文件中存在
TORSDOF的第二次出现,保留首次出现TORSDOF及之前的所有行,删除其后所有内容 - 处理后保留原文件名
示例文件
处理前:
$ cat FileWith2ndOccurance.txt ashu vishu jyoti TORSDOF Jatin Vishal Shivani TORSDOF Sushil Kiran
处理后:
after function run $ cat FileWith2ndOccurance.txt ashu vishu jyoti TORSDOF
实际文件样本(处理前)
REMARK Name = 17-DMAG.cdx REMARK 8 active torsions: REMARK status: ('A' for Active; 'I' for Inactive) REMARK 1 A between atoms: C_1 and N_8 REMARK 2 A between atoms: N_8 and C_9 REMARK 3 A between atoms: C_9 and C_10 REMARK 4 A between atoms: C_10 and N_11 REMARK 5 A between atoms: C_15 and O_17 REMARK 6 A between atoms: C_25 and O_28 REMARK 7 A between atoms: C_27 and O_33 REMARK 8 A between atoms: O_28 and C_29 REMARK x y z vdW Elec q Type REMARK _______ _______ _______ _____ _____ ______ ____ ROOT ATOM 1 C UNL 1 7.579 11.905 0.000 0.00 0.00 +0.000 C ATOM 2 C UNL 1 7.579 10.500 0.000 0.00 0.00 +0.000 C ATOM 30 O UNL 1 8.796 8.398 0.000 0.00 0.00 +0.000 OA ENDROOT BRANCH 21 31 ATOM 31 O UNL 1 13.701 7.068 0.000 0.00 0.00 +0.000 OA ATOM 32 C UNL 1 12.306 6.953 0.000 0.00 0.00 +0.000 C ENDBRANCH 41 42 ENDBRANCH 19 41 TORSDOF 8 REMARK Name = 17-DMAG.cdx REMARK 8 active torsions: REMARK status: ('A' for Active; 'I' for Inactive) REMARK 1 A between atoms: C_1 and N_8 REMARK 2 A between atoms: N_8 and C_9 REMARK x y z vdW Elec q Type REMARK _______ _______ _______ _____ _____ ______ ____ ROOT ATOM 1 CL UNL 1 0.000 11.656 0.000 0.00 0.00 +0.000 Cl ENDROOT TORSDOF 0
解决方案
以下是Python代码片段,可批量处理所有目标文件:
import os # 定义文件路径前缀和范围 file_prefix = "sade" start_num = 1 end_num = 7200 for num in range(start_num, end_num + 1): filename = f"{file_prefix}{num}.pdbqt" if not os.path.exists(filename): print(f"文件 {filename} 不存在,跳过") continue # 读取文件内容,记录首次TORSDOF及之前的行 keep_lines = [] torsdof_count = 0 with open(filename, 'r') as f: for line in f: keep_lines.append(line) if line.strip().startswith("TORSDOF"): torsdof_count += 1 # 当第二次出现时,停止读取并移除当前行(第二次的TORSDOF) if torsdof_count == 2: keep_lines.pop() break # 仅对存在多次TORSDOF的文件进行覆写 if torsdof_count >= 2: with open(filename, 'w') as f: f.writelines(keep_lines) print(f"已处理文件: {filename}") else: print(f"文件 {filename} 无需处理")
代码说明
- 遍历指定范围的所有pdbqt文件,跳过不存在的文件
- 逐行读取文件内容,计数
TORSDOF的出现次数 - 检测到第二次
TORSDOF时,立即停止读取,保留第一次出现及之前的所有行 - 仅修改存在多次
TORSDOF的文件,其他文件保持原样
注意事项
- 建议先备份部分文件测试代码效果,再批量处理所有文件
- 如果文件编码非默认utf-8,可在
open函数中添加encoding参数(如encoding='utf-8')
内容的提问来源于stack exchange,提问作者Jatin Kashyap
相关产品推荐
相关产品推荐

