You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

调整兰德指数(Adjusted Rand Score)原理、计算方法及Python实现问询

Understanding Adjusted Rand Score (ARS) & How to Use It in Python

Hey there! Let's break down exactly what Adjusted Rand Score is, how it calculates similarity between two clusterings, and how to use it with scikit-learn's metrics.adjusted_rand_score in Python.

What is Adjusted Rand Score?

First, let's ground this in the core idea you referenced:

兰德指数通过考虑所有样本对,统计在预测聚类与真实聚类中被分配到相同或不同簇的样本对数量,以此计算两个聚类结果间的相似度。

The raw Rand Score simply measures the share of sample pairs that are "consistent" between two clusterings: either both pairs are grouped together in both results, or both are split apart in both results.

But here's the catch: even a totally random clustering will have a non-zero Rand Score, especially with small datasets or specific cluster counts. The Adjusted Rand Score (ARS) fixes this by adjusting the raw score to eliminate bias from random chance, making it a far more reliable measure of true clustering similarity.

ARS ranges from -1 to 1:

  • 1 means the two clusterings are identical
  • 0 means the similarity is no better than random clustering
  • Negative values indicate the clustering is worse than random (rare in real-world use cases)

How is ARS Calculated?

Let’s walk through the math with a concrete example to make it tangible.

Key Definitions

We start with these values:

  • N: Total number of samples
  • true_labels: Ground-truth cluster assignments
  • pred_labels: Predicted cluster assignments

We’ll calculate four critical metrics based on all possible sample pairs:

  • a: Number of sample pairs that are in the same cluster in both true and predicted labels
  • sum_true_pairs: Total same-cluster pairs in the true labels
  • sum_pred_pairs: Total same-cluster pairs in the predicted labels
  • total_pairs: Total possible sample pairs (equal to N*(N-1)/2)

Step-by-Step Example

Let’s use a small dataset to demonstrate:

  • True clusters: [0, 0, 1, 1, 1] (2 samples in cluster 0, 3 in cluster 1)
  • Predicted clusters: [0, 0, 0, 1, 1] (3 samples in cluster 0, 2 in cluster 1)
  1. Calculate total sample pairs:
    total_pairs = 5*4/2 = 10

  2. Calculate sum_true_pairs:
    Sum of same-cluster pairs in the true labels:
    (2*1/2) + (3*2/2) = 1 + 3 = 4

  3. Calculate sum_pred_pairs:
    Sum of same-cluster pairs in the predicted labels:
    (3*2/2) + (2*1/2) = 3 + 1 = 4

  4. Calculate a:
    Count pairs that are grouped together in both clusterings using a contingency table of overlaps:

    • True cluster 0 & Predicted cluster 0: 2 samples → pairs = 2*1/2 = 1
    • True cluster 1 & Predicted cluster 1: 2 samples → pairs = 2*1/2 = 1
      Total a = 1 + 1 = 2
  5. Calculate expected a (random chance):
    This is the value of a we’d expect if clusters were assigned randomly:
    expected_a = (sum_true_pairs * sum_pred_pairs) / total_pairs = (4*4)/10 = 1.6

  6. Calculate maximum possible a:
    The highest feasible value of a given the true and predicted cluster sizes:
    max_a = (sum_true_pairs + sum_pred_pairs) / 2 = (4+4)/2 = 4

  7. Final ARS calculation:
    ARS = (a - expected_a) / (max_a - expected_a) = (2 - 1.6)/(4 - 1.6) ≈ 0.1667

Using metrics.adjusted_rand_score in Python

Scikit-learn handles all the complex math for you—here’s a quick implementation:

from sklearn import metrics

# Define ground-truth and predicted cluster labels
true_labels = [0, 0, 1, 1, 1]
pred_labels = [0, 0, 0, 1, 1]

# Compute Adjusted Rand Score
ars_score = metrics.adjusted_rand_score(true_labels, pred_labels)

# Print the result
print(f"Adjusted Rand Score: {ars_score:.4f}")

Running this will output:

Adjusted Rand Score: 0.1667

Quick Tips

  • ARS is symmetric: swapping true_labels and pred_labels gives the exact same score.
  • Label names don’t matter: if your true labels are [0,0,1,1] and predicted are [1,1,0,0], ARS will still be 1 (since the cluster groupings are identical).
  • It works seamlessly even if the true and predicted cluster counts don’t match.

内容的提问来源于stack exchange,提问作者bin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:40:47