You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

聚类性能评估:Python中是否有Rand Index可用包?附实现代码

Rand Index for Clustering Performance Evaluation

You're spot on—when using the Rand Index to evaluate clustering performance, Python doesn't have a direct out-of-the-box library function for the standard Rand Index. But as you mentioned, the Adjusted Rand Index is easy to access via sklearn.metrics.adjusted_rand_score(labels_true, labels_pred).

I recently needed the standard Rand Index for a project, so I wrote a working implementation that I'm sharing here for other developers to use:

import numpy as np
from itertools import combinations

def rand_index(labels_true, labels_pred):
    # Convert inputs to numpy arrays for streamlined processing
    labels_true = np.array(labels_true)
    labels_pred = np.array(labels_pred)
    
    # Generate all unique pairs of distinct samples
    sample_pairs = combinations(range(len(labels_true)), 2)
    
    concordant_pairs = 0
    discordant_pairs = 0
    
    for i, j in sample_pairs:
        # Check if the pair is grouped together in both label sets
        same_true_cluster = (labels_true[i] == labels_true[j])
        same_pred_cluster = (labels_pred[i] == labels_pred[j])
        
        if same_true_cluster == same_pred_cluster:
            concordant_pairs += 1
        else:
            discordant_pairs += 1
    
    # Calculate and return the Rand Index
    total_pairs = concordant_pairs + discordant_pairs
    return concordant_pairs / total_pairs if total_pairs != 0 else 1.0

You can use this function just like the sklearn adjusted variant: pass your ground truth labels (labels_true) and predicted clustering labels (labels_pred), and it will return the standard Rand Index score.

内容的提问来源于stack exchange,提问作者Hadij

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:24:18