You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

K-Means聚类算法准确率得分出现负值的原因咨询

Why You're Getting Negative "Accuracy" Scores with K-Means

Hey there! As someone who’s stumbled through my share of unsupervised learning mishaps early on, I totally get the confusion here. Let’s break down exactly what’s going on and how to fix it:

  • You’re using a supervised metric for an unsupervised algorithm
    Accuracy is made for supervised tasks (like classification) where you have known true labels to compare against predictions. But K-Means is unsupervised—it creates clusters without any knowledge of your true labels, and the cluster IDs it outputs (0, 1, 2, etc.) are totally arbitrary. For example, if your true labels are [0,0,1,1] but K-Means labels those same groups as [1,1,0,0], a raw accuracy score would be 0, not negative. That said, if you’re accidentally using a metric that can return negative values (like misinterpreting K-Means’ inertia as a "score," or using a custom loss function), that’s likely the culprit.

  • You might be miscalculating the score in code
    Double-check your implementation for these common mistakes:

    • Did you invert the score (e.g., returning -accuracy_score(true_labels, kmeans.labels_) instead of just the accuracy itself)?
    • Are you using a metric like the negative silhouette score or negative Calinski-Harabasz index by mistake? These metrics are meant to be maximized, so some workflows might flip their sign to turn them into "minimizable" loss values.
    • Did you mix up the order of true labels and predicted labels in the scoring function? Unlikely to cause negatives, but worth verifying.
  • Use the right metrics for K-Means instead of accuracy
    For evaluating unsupervised clustering, stick to metrics that don’t depend on matching arbitrary cluster IDs to true labels:

    • Adjusted Rand Index (ARI): Measures similarity between true labels and cluster assignments, adjusted for random chance. Ranges from -1 to 1 (1 = perfect match).
    • Silhouette Score: Quantifies how similar a data point is to its own cluster vs. other clusters. Ranges from -1 to 1 (higher = better clustering).
    • Calinski-Harabasz Index: Compares between-cluster variance to within-cluster variance. Higher values mean more distinct clusters.

Here’s a quick code swap to get you on the right track:
Instead of using accuracy:

from sklearn.metrics import accuracy_score
score = accuracy_score(true_labels, kmeans.labels_)

Try using ARI instead:

from sklearn.metrics import adjusted_rand_score
ari_score = adjusted_rand_score(true_labels, kmeans.labels_)

Don’t stress—this is an incredibly common pitfall when making the jump from supervised to unsupervised learning. You’re already asking the right questions, which is half the battle!

内容的提问来源于stack exchange,提问作者ashwin ram kumar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:34:56