层级聚类与K均值聚类实操问题:指定变量聚类及结果对比
Let’s walk through this step by step—starting with your data, prepping it for clustering, then implementing both hclust() and kmeans(), comparing results, and figuring out the optimal number of clusters. First, let’s load your sample data to make sure we’re on the same page:
df <- structure(list( id = c(111,111,111,112,112,112), se = c(1,2,3,1,2,3), t1 = c(1, 2, 1, 1,1,3), t2 = c(1, 2, 2, 1,1,4), t3 = c(1, 0, 0, 0,2,1), t4 = c(2, 5, 7, 7,1,2), t5 = c(1, 0, 1, 1,1,1), t6 = c(1, 1, 1, 1,1,1), t7 = c(1, 1, 1 ,1,1,1), t8=c(0,0,0,0,0,0) ), row.names = c(NA, 6L), class = "data.frame")
1. Data Preparation: Focus on the t2 Metric
Your data has multiple rows per id (one for each se). First, we need to decide what to cluster:
- Option 1: Cluster individual IDs (using their
t2values across allselevels) - Option 2: Cluster individual rows (each
seentry for eachid)
Let’s cover both, starting with the more common use case of clustering IDs.
Option 1: Reshape Data for ID-Level Clustering
We’ll pivot the data so each row is an ID, with columns representing t2 values for each se:
library(tidyr) library(dplyr) # Reshape to get one row per ID, with t2 values for each se clust_data <- df %>% select(id, se, t2) %>% pivot_wider(names_from = se, values_from = t2, names_prefix = "t2_se_") %>% column_to_rownames("id") # Set ID as row names for easier clustering # Check the result clust_data
This gives us a clean dataset where each row is an ID, and columns are their t2 values across se=1,2,3.
Option 2: Use Individual Rows for Clustering
If you want to cluster the 6 rows directly using only t2, just extract the t2 column:
row_clust_data <- df %>% select(t2)
For the rest of this guide, we’ll use the ID-level clustered data (Option 1), but you can adapt the code to Option 2 easily.
2. Hierarchical Clustering with hclust()
Step 2.1: Calculate Distance Between Observations
First, compute a distance matrix (Euclidean distance is the default, but you can use Manhattan or others):
dist_matrix <- dist(clust_data, method = "euclidean")
Step 2.2: Run Hierarchical Clustering
Choose a linkage method—ward.D2 is great for minimizing within-cluster variance, but you can also try complete or single:
hier_clust <- hclust(dist_matrix, method = "ward.D2")
Step 2.3: Visualize the Dendrogram
Plot the dendrogram to see how observations group together:
plot(hier_clust, main = "Hierarchical Clustering Dendrogram (Based on t2)", xlab = "ID", ylab = "Distance", cex = 0.8)
Step 2.4: Find the Optimal Number of Clusters
Use these methods to pick the best k:
- Elbow Method: Look for where the curve flattens (using the
factoextrapackage):
library(factoextra) fviz_nbclust(clust_data, FUN = hcut, method = "wss") + ggtitle("Elbow Method for Optimal Clusters")
- Silhouette Method: Higher scores mean better cluster separation:
fviz_nbclust(clust_data, FUN = hcut, method = "silhouette") + ggtitle("Silhouette Method for Optimal Clusters")
Step 2.5: Assign Cluster Labels
Once you choose k (e.g., k=2), assign clusters to each ID:
k_optimal <- 2 hier_clusters <- cutree(hier_clust, k = k_optimal) # Add cluster labels to your data clust_data$hier_cluster <- hier_clusters
3. K-Means Clustering with kmeans()
Step 3.1: Run K-Means
Use the same optimal k from hierarchical clustering (or find it separately):
# Run k-means with 2 clusters, 20 initial starts to avoid local minima kmeans_clust <- kmeans(clust_data %>% select(-hier_cluster), centers = k_optimal, nstart = 20)
Step 3.2: Analyze Results
Check key outputs from k-means:
# Cluster centers (average t2 values for each cluster) kmeans_clust$centers # Cluster assignments for each ID kmeans_clust$cluster # Total within-cluster variance (lower = better) kmeans_clust$tot.withinss
Step 3.3: Assign K-Means Labels
Add the labels to your dataset:
clust_data$kmeans_cluster <- kmeans_clust$cluster
4. Compare Hierarchical & K-Means Results
See how the two methods assign clusters to IDs:
# View side-by-side cluster assignments clust_data # Create a contingency table to compare table(clust_data$hier_cluster, clust_data$kmeans_cluster) # Visualize the comparison library(ggplot2) clust_data %>% rownames_to_column("id") %>% ggplot(aes(x = hier_cluster, y = kmeans_cluster, color = id)) + geom_point(size = 4) + labs(title = "Hierarchical vs K-Means Cluster Assignments", x = "Hierarchical Cluster", y = "K-Means Cluster")
Quick Tips for Your Workflow
- Data Shape: Always ensure each row is an observation (ID, row, etc.) and columns are features (like
t2values). - Method Selection: For hierarchical clustering, experiment with different linkage methods to see which fits your data best.
- Optimal K: Don’t rely on one method—combine elbow/silhouette results with your domain knowledge to pick k.
内容的提问来源于stack exchange,提问作者Rstudent

