关于层次聚类(Hierarchical Clustering)与K-means聚类的聚类数疑惑
Great question—this is one of the most common sticking points when switching between K-means and hierarchical clustering, so you’re not alone in this confusion!
Let’s break this down clearly:
- First, it’s true that hierarchical clustering doesn’t spit out a fixed number of clusters by default like K-means does. Instead, it builds a hierarchical structure (dendrogram) that shows how every sample connects to others through merging (agglomerative) or splitting (divisive) steps. But this flexibility is actually a feature, not a limitation—you just need a way to pick the right number of clusters from that structure.
How to determine the number of clusters from hierarchical clustering
Here are the most practical methods:
Dendrogram Cutting
The dendrogram is your best visual tool. Each vertical line represents a cluster merge/split, and the height of the horizontal lines shows the distance (or dissimilarity) at which the merge happened. To pick a cluster count:- Look for large gaps between horizontal lines. These gaps indicate that merging those clusters would require a big jump in dissimilarity—meaning those groups are naturally distinct. Cut the dendrogram horizontally at that gap level, and the number of vertical lines below the cut is your cluster count.
- For example, if your dendrogram has a huge gap between height 4 and 6, cutting at height 5 might give you 3 distinct clusters that make intuitive sense for your data.
Statistical Metrics
If you want a more objective approach, use the same metrics you might use for K-means:- Elbow Method: Calculate the total within-cluster sum of squares (WCSS) for different cluster counts (by cutting the dendrogram at different heights) and plot it. The "elbow" point where the WCSS stops dropping sharply is your optimal k.
- Silhouette Score: Compute the average silhouette coefficient for each possible k. The k with the highest average score indicates clusters that are both tight internally and distinct from each other.
- Calinski-Harabasz Index: This measures the ratio of between-cluster variance to within-cluster variance. Higher values mean better clustering—pick the k with the highest index value.
Directly Specify k in Implementation
Most modern ML libraries let you skip the dendrogram step if you already have a guess for k. For example, in scikit-learn’sAgglomerativeClustering, you can set then_clustersparameter directly, and the algorithm will automatically cut the hierarchy to give you exactly that many clusters.
Quick comparison to K-means
K-means forces you to choose k upfront, which can be a downside if you don’t have domain knowledge about your data. Hierarchical clustering lets you explore the full structure first, then decide on k—perfect for exploratory analysis. On the flip side, K-means is faster for large datasets, while hierarchical clustering can be slower but gives you more interpretability via the dendrogram.
内容的提问来源于stack exchange,提问作者Goktug

