SMILES字符串分支计数功能异常求助:代码修正需求
SMILES分支计数功能修正
问题说明
需要实现SMILES字符串的分支计数功能,规则如下:
- 以括号表示分支,分支内的原子使用同一计数编号,嵌套括号暂不单独处理(如
CC(CC(C))需输出[0,0,1,1,1]) - 同一主链原子后的多个分支,编号依次递增(如
C(N)(N)需输出[0,1,2]) - 回到非分支区域时计数重置为0
- 忽略SMILES中的括号和环编号数字,仅统计原子的分支计数
实际输出
SMILES branch_count 0 C(N)(N)CC(N)C [0, 0, 1, 0, 0, 1, 0, 0, 0, 0, 1, 0, 0] 1 CCC [0, 0, 0] 2 C1CC1 [0, 0, 0, 0, 0] 3 C1CC1(C)C [0, 0, 0, 0, 0, 0, 1, 0, 0] 4 CC(C)C [0, 0, 0, 1, 0, 0]
期望输出
SMILES branch_count 0 C(N)(N)CC(N)C [0, 1, 2, 0, 0, 1, 0] 1 CCC [0, 0, 0] 2 C1CC1 [0, 0, 0] 3 C1CC1(C)C [0, 0, 0, 1, 0] 4 CC(C)C [0, 0, 1, 0]
当前代码
import pandas as pd import numpy as np from rdkit import Chem def get_branch_count(smile): # Initialize variables branch_count = [0] * len(smile) bracket_count = 0 current_count = 0 # Loop through each character in the smile for i, c in enumerate(smile): # If the character is an open bracket, increment bracket count if c == "(": bracket_count += 1 # If the character is a close bracket, decrement bracket count elif c == ")": bracket_count -= 1 # If there are no more open brackets after this one, reset current count if bracket_count == 0: current_count = 0 # If the character is not a bracket, update the current count else: if bracket_count > 0: # If the previous character was also a bracket, don't increment the count if smile[i-1] != ")": current_count += 1 else: current_count = 0 branch_count[i] = current_count return branch_count def collect_branch_count(smile_list): rows = [] for smile in smile_list: branch_count = get_branch_count(smile) data = {"branch_count": branch_count} row = {"SMILES": smile} for key, value in data.items(): row[key] = value rows.append(row) df = pd.DataFrame(rows) return df smile_list = ["C(N)(N)CC(N)C", "CCC", "C1CC1", "C1CC1(C)C", "CC(C)C"] df = collect_branch_count(smile_list) print(df)
修正后的代码
import pandas as pd def get_branch_count(smile): branch_count = [] bracket_level = 0 current_main_branch_count = 0 current_branch_num = 0 in_branch = False for c in smile: if c == "(": bracket_level += 1 # 进入最外层分支,更新当前分支编号 if bracket_level == 1: current_main_branch_count += 1 current_branch_num = current_main_branch_count in_branch = True elif c == ")": bracket_level -= 1 # 回到主链,标记退出分支 if bracket_level == 0: in_branch = False elif c.isdigit(): # 忽略环编号数字 continue else: # 处理原子字符 if in_branch: branch_count.append(current_branch_num) else: branch_count.append(0) # 新的主链原子,重置其分支计数 current_main_branch_count = 0 return branch_count def collect_branch_count(smile_list): rows = [] for smile in smile_list: branch_count = get_branch_count(smile) rows.append({ "SMILES": smile, "branch_count": branch_count }) return pd.DataFrame(rows) smile_list = ["C(N)(N)CC(N)C", "CCC", "C1CC1", "C1CC1(C)C", "CC(C)C"] df = collect_branch_count(smile_list) print(df)
修正说明
- 过滤无关字符:遍历SMILES时跳过括号和环编号数字,仅保留原子字符,确保结果数组长度与原子数量一致
- 分支计数逻辑优化:
- 跟踪括号层级,仅在进入最外层分支时更新分支编号
- 记录当前主链原子的分支数量,同一主链原子后的多个分支依次递增编号
- 分支内的所有原子使用同一编号,嵌套括号不改变当前分支编号(符合需求中“嵌套括号后续单独处理”的要求)
- 主链原子重置:遇到新的主链原子时,重置其分支计数,确保不同主链原子的分支编号独立计数
内容的提问来源于stack exchange,提问作者YZman
相关产品推荐
相关产品推荐

