如何确定Python中Affinity Propagation聚类图像簇的样本点?
Hey there! I’ve worked with Affinity Propagation (AP) for image clustering before, so let me walk you through exactly how to find those exemplar images for each cluster—especially if you’re using scikit-learn’s implementation (which is the go-to for Python).
核心逻辑:利用AP模型的内置属性
When you train an Affinity Propagation model, it automatically tracks which samples become exemplars. The key is two attributes the model outputs after fitting:
cluster_centers_indices_: This is an array of integers where each value is the index of your original image dataset corresponding to a cluster’s exemplar. For example, if this array returns[3, 10, 17], the 4th, 11th, and 18th images (since Python uses 0-indexing) are your three cluster representatives.labels_: This array has the same length as your input dataset, where each element is the cluster ID assigned to that image. You can use this to map every image back to its cluster’s exemplar.
实战代码示例
Let’s say you’ve already extracted features from your images (like using a pre-trained CNN to get embedding vectors, stored in image_features). Here’s how to tie it all together:
from sklearn.cluster import AffinityPropagation # 假设image_features是形状为(n_images, n_features)的特征数组 ap_model = AffinityPropagation(random_state=42, preference=-50) # preference控制簇数量,按需调整 ap_model.fit(image_features) # 1. 获取所有簇的代表图像索引 exemplar_indices = ap_model.cluster_centers_indices_ print("所有簇的代表图像索引:", exemplar_indices) # 2. 获取每个图像对应的簇标签 image_cluster_labels = ap_model.labels_ # 3. 遍历每个簇,匹配对应的代表图像 unique_clusters = set(image_cluster_labels) for cluster_id in unique_clusters: # 找到当前簇的所有图像索引 cluster_image_ids = [idx for idx, label in enumerate(image_cluster_labels) if label == cluster_id] # 当前簇的代表图像索引 current_exemplar_idx = exemplar_indices[cluster_id] print(f"簇 {cluster_id} 的代表图像是第 {current_exemplar_idx} 张图,包含的图像索引:{cluster_image_ids}")
如果是自定义实现AP的情况
If you’re rolling your own Affinity Propagation algorithm instead of using scikit-learn, remember that an exemplar is a sample where the sum of its responsibility and availability values is the highest in its cluster. You’ll need to track these two metrics during the algorithm’s iteration, and once convergence is reached, identify samples where this sum is maximized for each cluster.
小提示
- The
preferenceparameter in scikit-learn’s AP controls how many clusters you get. Higher preference values lead to more clusters (and more exemplars), lower values lead to fewer. Play around with this to match your image clustering needs.
内容的提问来源于stack exchange,提问作者IkramAc

