聚类性能评估:Python中是否有Rand Index可用包?附实现代码
Rand Index for Clustering Performance Evaluation
You're spot on—when using the Rand Index to evaluate clustering performance, Python doesn't have a direct out-of-the-box library function for the standard Rand Index. But as you mentioned, the Adjusted Rand Index is easy to access via sklearn.metrics.adjusted_rand_score(labels_true, labels_pred).
I recently needed the standard Rand Index for a project, so I wrote a working implementation that I'm sharing here for other developers to use:
import numpy as np from itertools import combinations def rand_index(labels_true, labels_pred): # Convert inputs to numpy arrays for streamlined processing labels_true = np.array(labels_true) labels_pred = np.array(labels_pred) # Generate all unique pairs of distinct samples sample_pairs = combinations(range(len(labels_true)), 2) concordant_pairs = 0 discordant_pairs = 0 for i, j in sample_pairs: # Check if the pair is grouped together in both label sets same_true_cluster = (labels_true[i] == labels_true[j]) same_pred_cluster = (labels_pred[i] == labels_pred[j]) if same_true_cluster == same_pred_cluster: concordant_pairs += 1 else: discordant_pairs += 1 # Calculate and return the Rand Index total_pairs = concordant_pairs + discordant_pairs return concordant_pairs / total_pairs if total_pairs != 0 else 1.0
You can use this function just like the sklearn adjusted variant: pass your ground truth labels (labels_true) and predicted clustering labels (labels_pred), and it will return the standard Rand Index score.
内容的提问来源于stack exchange,提问作者Hadij
相关产品推荐
相关产品推荐

