You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

SDF文件化合物命名批量添加计数后缀异常排查求助

问题:给SDF文件中重复的化合物名称添加计数编号

原始SDF文件格式

$$$$
compound1
#lots of text


$$$$
compound1
#lots of text

$$$$
compound2
#lots of text

$$$$
compound2
#lots of text

预期处理结果

$$$$
compound1_1
#lots of text


$$$$
compound1_2
#lots of text

$$$$
compound2_1
#lots of text

$$$$
compound2_2
#lots of text

当前使用的Python代码

import sys
import re
import os
file = sys.argv[1]
file2 = file + '_fixed'
name2 = 'compound'
with open(file2,'w') as new_file:
    with open(file) as fp:
        # read all lines in a list
        for line in fp:
        # check if string present on a current line
            if "$$$$" in line:
                new_file.write(line)
                name = next(fp)
                if name2 != name:
                    name2 = name
                    j = 0
                    new_name = name + '_' + str(j)
                    new_file.write(new_name)
                else:
                    j = j + 1
                    new_name = name + '_' + str(j)
                    new_file.write(new_name)
            else:
                new_file.write(line)

实际错误输出

$$$$
compound1
_1 #lots of text


$$$$
compound1
_2 #lots of text

$$$$
compound2
_1 #lots of text

$$$$
compound2
_2 #lots of text

问题分析与解决思路

核心问题

  1. 换行符未处理:读取的化合物名称name包含行尾的\n,直接拼接编号会导致_数字被写到新一行,这是格式错乱的主要原因。
  2. 计数逻辑错误:首次遇到新化合物时初始化j=0,直接写入_0不符合预期的从1开始计数;同时用单变量j记录计数,切换化合物时容易出现逻辑混乱。
  3. 迭代器使用风险:next(fp)在循环中可能会跳过部分行,导致内容丢失或处理异常。

修正方案

用字典单独记录每个化合物的计数,同时处理换行符,改用更稳妥的逐行读取方式:

import sys

file = sys.argv[1]
file2 = file + '_fixed'
# 字典存储每个化合物的当前计数,初始为0
compound_counts = {}

with open(file2, 'w') as new_file:
    with open(file, 'r') as fp:
        current_line = fp.readline()
        while current_line:
            if "$$$$" in current_line:
                new_file.write(current_line)
                # 读取化合物名称,去除行尾换行符
                compound_name = fp.readline().rstrip('\n')
                # 更新该化合物的计数:如果不存在则从1开始,否则+1
                compound_counts[compound_name] = compound_counts.get(compound_name, 0) + 1
                # 拼接新名称并添加换行符,写入文件
                new_name = f"{compound_name}_{compound_counts[compound_name]}\n"
                new_file.write(new_name)
                # 读取下一行,继续循环
                current_line = fp.readline()
            else:
                new_file.write(current_line)
                current_line = fp.readline()

修正说明

  • 用rstrip('\n')移除化合物名称的换行符,确保编号和名称在同一行。
  • 字典compound_counts独立跟踪每个化合物的计数,逻辑清晰,不会因化合物切换出错。
  • 计数从1开始,完全匹配预期结果。
  • 改用while循环配合readline(),避免next(fp)带来的迭代器异常,保证每一行都被正确处理。

内容的提问来源于stack exchange,提问作者Phung Hien Le

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 12:07:37