残基对数据集的K-means聚类实现及方法选型咨询
残基对数据集的聚类与可视化问题
我有一个包含R1(残基1)、R2(残基2)、**Score(得分)**三列的CSV数据集,需求如下:
- 先为每个残基对分配唯一编号
- 按得分区间(>0.8、0.3-0.8、<0.3)完成聚类可视化
我尝试过K-means聚类但未得到预期结果,也用Pyviz做了网络可视化(仅用于观察残基连接情况),现在想咨询:
- 这个数据集是否适合做聚类分析?
- 应该选择K-means还是其他聚类技术?
已尝试的K-means聚类代码
import pandas as pd import matplotlib.pyplot as plt from sklearn.cluster import KMeans import numpy as np csv_file_path = 'contact.csv' data = pd.read_csv(csv_file_path) pairs = data[['R1', 'R2']].values.tolist() mi_values = data['Score'].values.tolist() num_clusters = 3 mi_values_array = np.array(mi_values).reshape(-1, 1) kmeans = KMeans(n_clusters=num_clusters, random_state=0).fit(mi_values_array) data['cluster'] = kmeans.labels_ colors = ['b', 'g', 'r', 'c', 'm', 'y', 'k'] # Add more colors if needed plt.figure(figsize=(8, 6)) for cluster_id in range(num_clusters): cluster_residues = data[data['cluster'] == cluster_id][['R1', 'R2']].values.tolist() x = [pair[0] for pair in cluster_residues] y = [pair[1] for pair in cluster_residues] plt.scatter(x, y, label=f'Cluster {cluster_id + 1}', c=colors[cluster_id]) plt.xlabel('Residue 1') plt.ylabel('Residue 2') plt.title('Residue Pairs Clustering') plt.legend() plt.show()
已尝试的Pyviz网络可视化代码
import networkx as nx from pyvis.network import Network import pandas as pd from IPython.core.display import Image from PIL import Image import matplotlib.pyplot as plt filtered_df = data[data['Score'] > 0] # Extract unique residue pairs from the filtered DataFrame residue_pairs = set(zip(filtered_df['R1'], filtered_df['R2'])) # Create an empty undirected graph G = nx.Graph() # Add the residue pairs as nodes to the graph G.add_nodes_from(residue_pairs) # Create an interactive network visualization net = Network(notebook=True) # Create a dictionary for node colors node_colors = {} # Add nodes and edges to the interactive network for node in G.nodes(): res1_label = node[0] res2_label = node[1] mi_score = 0.0 # Initialize mi_score to a default value # Find the corresponding row in the filtered DataFrame # and get the MI score if (res1_label, res2_label) in zip(filtered_df['R1'], filtered_df['R2']): mi_score = filtered_df[(filtered_df['R1'] == res1_label) & (filtered_df['R2'] == res2_label)]['Score'].values[0] # Define color based on the MI score conditions if mi_score >= 0.8: node_colors[node] = 'red' elif 0.3 <= mi_score < 0.8: node_colors[node] = 'blue' else: node_colors[node] = 'green' # Add nodes with color to the interactive network net.add_node(res1_label, label=f'Res1: {res1_label}', color=node_colors[node]) net.add_node(res2_label, label=f'Res2: {res2_label}', color=node_colors[node]) net.add_edge(res1_label, res2_label) # Show the interactive graph net.show_buttons(filter_=['nodes', 'edges', 'physics']) net.show('interactive_graph.html')
问题解答
1. 数据集是否适合聚类?
完全适合。你的核心需求是基于Score的区间对残基对分组,本质是基于数值特征的分组聚类,这类场景非常适合聚类分析,尤其是当你需要将残基对按得分强度区分开时。
2. K-means是否合适?为什么之前没得到预期结果?
K-means在这里不是最优选择,原因如下:
- K-means是基于距离的无监督聚类算法,它会自动将数据分成密度相近的簇,但你已经明确了固定的得分区间规则,属于有监督的分组/分类,而非无监督聚类。
- 你之前的代码仅用Score单特征做K-means,K-means会根据数据分布自动划分簇边界,不一定会贴合你设定的>0.8、0.3-0.8、<0.3这三个区间,这就是你没得到预期结果的核心原因。
更适合的方案
方案一:直接按规则分组(最贴合需求)
既然你已经明确了得分区间,不需要无监督聚类,直接用条件判断给每个残基对分配分组标签即可,代码示例:
import pandas as pd import matplotlib.pyplot as plt # 读取数据 data = pd.read_csv('contact.csv') # 1. 为残基对分配唯一编号 data['pair_id'] = range(1, len(data)+1) # 2. 按得分区间分组 def assign_group(score): if score > 0.8: return 'High (>0.8)' elif 0.3 <= score <= 0.8: return 'Medium (0.3-0.8)' else: return 'Low (<0.3)' data['group'] = data['Score'].apply(assign_group) # 可视化 colors = {'High (>0.8)': 'red', 'Medium (0.3-0.8)': 'blue', 'Low (<0.3)': 'green'} plt.figure(figsize=(8,6)) for group, color in colors.items(): subset = data[data['group'] == group] plt.scatter(subset['R1'], subset['R2'], label=group, c=color, alpha=0.6) plt.xlabel('残基1') plt.ylabel('残基2') plt.title('残基对按得分区间分组可视化') plt.legend() plt.show()
方案二:若坚持用聚类算法(适合不确定区间的场景)
如果你希望算法自动发现得分的聚类模式,而非用固定区间,可以考虑:
- DBSCAN:适合基于密度的聚类,能识别出得分分布中的密集簇,比如高得分的簇、中等得分的簇等。
- 层次聚类:可以生成聚类树,帮助你观察得分的层级分布,再手动切割成你需要的3个簇。
以DBSCAN为例的代码示例:
import pandas as pd import matplotlib.pyplot as plt from sklearn.cluster import DBSCAN import numpy as np data = pd.read_csv('contact.csv') data['pair_id'] = range(1, len(data)+1) # 用Score作为特征 X = np.array(data['Score']).reshape(-1,1) # 调整eps和min_samples参数适配你的数据分布 dbscan = DBSCAN(eps=0.1, min_samples=5).fit(X) data['cluster'] = dbscan.labels_ # 可视化 plt.figure(figsize=(8,6)) unique_clusters = data['cluster'].unique() colors = plt.cm.get_cmap('tab10', len(unique_clusters)) for idx, cluster in enumerate(unique_clusters): subset = data[data['cluster'] == cluster] plt.scatter(subset['R1'], subset['R2'], label=f'Cluster {cluster}', c=colors(idx), alpha=0.6) plt.xlabel('残基1') plt.ylabel('残基2') plt.title('DBSCAN残基对聚类可视化') plt.legend() plt.show()
关于Pyviz可视化的优化
你之前的Pyviz代码存在一个问题:误将残基对作为图的节点,正确的做法应该是将残基作为节点,残基对作为边,边的颜色/粗细由Score决定,这样才能准确展示残基之间的连接强度:
import pandas as pd from pyvis.network import Network data = pd.read_csv('contact.csv') filtered_df = data[data['Score'] > 0] net = Network(notebook=True, height='800px', width='100%') # 添加残基节点(去重) all_residues = set(filtered_df['R1'].tolist() + filtered_df['R2'].tolist()) for res in all_residues: net.add_node(res, label=f'残基{res}') # 添加边,根据得分设置颜色和粗细 for _, row in filtered_df.iterrows(): res1, res2, score = row['R1'], row['R2'], row['Score'] if score > 0.8: color = 'red' width = 3 elif 0.3 <= score <=0.8: color = 'blue' width = 2 else: color = 'green' width = 1 net.add_edge(res1, res2, color=color, width=width, title=f'Score: {score:.2f}') net.show_buttons(filter_=['nodes', 'edges', 'physics']) net.show('optimized_interactive_graph.html')
内容的提问来源于stack exchange,提问作者tanya singhal
相关产品推荐
相关产品推荐

