如何基于uniq-compounds提取指定文件夹out.pdbqt中的化合物名称
嘿,我来给你分享两个实用的方案,帮你快速完成这个筛选和提取的任务——不管你更习惯用Shell脚本还是Python,都能轻松搞定:
Bash 脚本方案
如果你平时经常用命令行,这个脚本会很顺手,直接在终端就能跑:
#!/bin/bash # 假设uniq-compounds文件和sidra目录都在当前工作目录下 while read -r ligand_dir; do # 拼接出目标out.pdbqt的完整路径 pdbqt_file="./sidra/${ligand_dir}/out.pdbqt" # 先检查文件是否存在,避免报错 if [ -f "$pdbqt_file" ]; then # 提取第3行,再截取REMARK里"Name = "后面的化合物名称 compound_name=$(sed -n '3p' "$pdbqt_file" | sed 's/^REMARK.*Name = //') # 将结果写入compound_names.txt,格式为「文件夹名: 化合物名称」 echo "${ligand_dir}: ${compound_name}" >> compound_names.txt else # 输出警告,提示找不到文件,方便排查 echo "⚠️ 警告:$pdbqt_file 文件不存在" >&2 fi done < uniq-compounds
脚本说明:
read -r ligand_dir逐行读取uniq-compounds里的每个文件夹名sed -n '3p'精准提取out.pdbqt的第3行内容- 第二个
sed命令负责去掉行首的REMARK前缀,只保留Name =后面的化合物名称 - 最终结果会追加到
compound_names.txt文件中,同时终端会打印缺失文件的警告
Python 脚本方案
如果你更熟悉Python,这个脚本可读性更强,也更容易调整逻辑:
# 读取uniq-compounds里的文件夹列表,去掉空行和多余空格 with open('uniq-compounds', 'r') as f: ligand_dirs = [line.strip() for line in f if line.strip()] compound_list = [] for ligand_dir in ligand_dirs: pdbqt_path = f"./sidra/{ligand_dir}/out.pdbqt" try: with open(pdbqt_path, 'r') as pdbqt_file: lines = pdbqt_file.readlines() # 确保文件至少有3行 if len(lines) >= 3: third_line = lines[2] # Python索引从0开始,第3行对应索引2 if 'Name = ' in third_line: # 截取"Name = "后面的内容并去除首尾空格 compound_name = third_line.split('Name = ')[1].strip() compound_list.append(f"{ligand_dir}: {compound_name}") else: print(f"⚠️ 警告:{pdbqt_path}的第3行没有找到Name字段") else: print(f"⚠️ 警告:{pdbqt_path}的行数不足3行") except FileNotFoundError: print(f"⚠️ 警告:{pdbqt_path}文件不存在") # 将结果写入compound_names.txt with open('compound_names.txt', 'w') as out_file: out_file.write('\n'.join(compound_list))
脚本说明:
- 用列表推导式读取并整理文件夹列表,自动过滤空行
- 用
try-except捕获文件不存在的异常,避免脚本中途崩溃 - 先判断文件行数是否足够,再检查是否有Name字段,逻辑更严谨
- 最终结果同样保存到
compound_names.txt中,同时在终端输出异常提示
小提示:
- 如果
uniq-compounds或sidra不在当前目录,记得修改脚本里的路径(比如把./sidra改成绝对路径/home/xxx/sidra) - 可以根据需求调整输出格式,比如只保留化合物名称,或者把结果导入到Excel里
- 两种方案都会自动跳过无效文件,不会因为个别缺失文件影响整体任务
内容的提问来源于stack exchange,提问作者Sidra Rafi
相关产品推荐
相关产品推荐

