You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

候选药物分子相似性匹配最优算法及加权方案咨询

药物分子相似性计算:最优算法、权重赋值与可视化方案

一、最优相似性度量算法选择

不同度量算法的结果差异源于对属性差异的敏感逻辑不同,针对药物分子多属性场景,推荐方案如下:

  • 优先选用加权欧氏距离:适配所有连续/计数型属性,可处理量纲差异(比如表面积、氢键供体数量的量级差),同时支持灵活权重赋值,是药物研发中相似性计算的主流选择。
  • 备选:曼哈顿距离:若属性中离散计数型(如D、A)占比高,该算法对异常值鲁棒性更好,但灵活性弱于加权欧氏。
  • 避坑提示:
    • 余弦相似度:仅关注向量方向,忽略属性绝对值差异,不适合药物分子的定量属性对比(如分子量、表面积的绝对值直接影响成药性)。
    • 切比雪夫距离:仅聚焦差异最大的单个属性,会丢失其他多维度特征的贡献,不符合药物多属性相似性的核心需求。

二、属性权重赋值方法

1. 领域知识驱动赋值

结合药物研发的生物学意义直接分配权重,例如:

  • 核心指标**蛋白结合亲和力(Vina)**赋予最高权重(如0.3);
  • 影响成药性的极性表面积(PSA)、**氢键供体/受体(D/A)**各赋0.15;
  • 分子量(Mw)、**表面积(Su)**等结构属性各赋0.1。

2. 数据统计驱动赋值

无明确领域知识时,基于数据特征计算权重:

  • 归一化方差权重:对属性做Z-score归一化后,用属性方差作为权重(方差越大,区分度越高,权重越高);
  • 互信息权重:计算每个属性与目标药物活性标签的互信息,互信息越高权重越高。

3. 在scipy中实现加权距离

scipy.spatial.distance.cdist未直接提供加权欧氏距离,但可通过标准化欧氏距离转换实现:

import numpy as np
from scipy.spatial.distance import cdist

# 假设data是候选分子属性矩阵(已剔除Mol列),target是目标药物属性向量
# 先做min-max归一化消除量纲影响
data_normalized = (data - data.min(axis=0)) / (data.max(axis=0) - data.min(axis=0))
target_normalized = (target - data.min(axis=0)) / (data.max(axis=0) - data.min(axis=0))

# 定义权重数组,顺序对应Su, Vol, PSA, Ov, D, A, Mw, Vina
weights = np.array([0.1, 0.1, 0.15, 0.1, 0.15, 0.15, 0.1, 0.3])

# 用seuclidean实现加权欧氏距离,V参数传入1/weights²转换权重
weighted_distances = cdist(
    data_normalized, 
    target_normalized.reshape(1, -1), 
    metric='seuclidean', 
    V=1/weights**2
).flatten()

# 获取最相似分子索引
most_similar_idx = np.argmin(weighted_distances)

注:标准化欧氏距离的V参数是方差数组,用1/weights²等价于为每个维度赋予对应权重。

三、相似性结果可视化

1. 雷达图:多属性对比

直观展示候选分子与目标药物的各维度差异:

import matplotlib.pyplot as plt
import pandas as pd

# 假设df包含所有分子数据,最后一行为目标药物
target = df.iloc[-1]
candidates = df.iloc[:-1]
top2_similar = candidates.loc[np.argsort(weighted_distances)[:2]]

attributes = ['Su', 'Vol', 'PSA', 'Ov', 'D', 'A', 'Mw', 'Vina']
num_attr = len(attributes)

# 设置雷达图角度
angles = np.linspace(0, 2*np.pi, num_attr, endpoint=False)
angles = np.concatenate((angles, [angles[0]]))

# 归一化属性数据
df_normalized = (df[attributes] - df[attributes].min()) / (df[attributes].max() - df[attributes].min())
target_norm = df_normalized.iloc[-1].values
target_norm = np.concatenate((target_norm, [target_norm[0]]))

# 绘制雷达图
fig, ax = plt.subplots(figsize=(8,8), subplot_kw={'polar': True})
ax.plot(angles, target_norm, 'o-', linewidth=2, label='Target Drug')
for idx, row in top2_similar.iterrows():
    row_norm = df_normalized.loc[idx].values
    row_norm = np.concatenate((row_norm, [row_norm[0]]))
    ax.plot(angles, row_norm, 'o-', linewidth=2, label=f'Molecule {row["Mol"]}')

ax.set_xticks(angles[:-1])
ax.set_xticklabels(attributes)
ax.set_title('Drug Molecule Attribute Comparison', y=1.1)
ax.legend(loc='upper right', bbox_to_anchor=(0.1, 0.1))
plt.show()

2. 柱状图:相似度排序

展示所有候选分子与目标药物的相似度(距离越小相似度越高,转换为1/(1+距离)将值映射到[0,1]区间):

# 计算相似度得分
similarity_scores = 1 / (1 + weighted_distances)

# 按相似度降序排序
sorted_indices = np.argsort(similarity_scores)[::-1]
sorted_mols = candidates['Mol'].iloc[sorted_indices].values
sorted_scores = similarity_scores[sorted_indices]

# 绘制柱状图
plt.figure(figsize=(12,6))
plt.bar(sorted_mols, sorted_scores, color='#4287f5')
plt.axhline(y=similarity_scores[most_similar_idx], color='r', linestyle='--', label='Top Similar')
plt.title('Candidate Molecules Similarity to Target Drug')
plt.xlabel('Molecule ID')
plt.ylabel('Similarity Score')
plt.xticks(rotation=45)
plt.legend()
plt.tight_layout()
plt.show()

3. 聚类热图:分子分组

通过层次聚类展示所有分子的属性相似性分组:

import seaborn as sns

# 归一化属性矩阵
norm_data = df_normalized.values

# 绘制聚类热图
plt.figure(figsize=(10,8))
sns.clustermap(
    norm_data, 
    row_labels=df['Mol'], 
    col_labels=attributes, 
    cmap='coolwarm', 
    figsize=(12,10)
)
plt.title('Drug Molecule Attribute Clustering')
plt.show()

内容的提问来源于stack exchange,提问作者Muzeita

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 06:45:33