You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

层级聚类与K均值聚类实操问题:指定变量聚类及结果对比

Step-by-Step Guide to Hierarchical & K-Means Clustering in R (Using Your Data)

Let’s walk through this step by step—starting with your data, prepping it for clustering, then implementing both hclust() and kmeans(), comparing results, and figuring out the optimal number of clusters. First, let’s load your sample data to make sure we’re on the same page:

df <- structure(list(
  id = c(111,111,111,112,112,112), 
  se = c(1,2,3,1,2,3),
  t1 = c(1, 2, 1, 1,1,3), 
  t2 = c(1, 2, 2, 1,1,4), 
  t3 = c(1, 0, 0, 0,2,1), 
  t4 = c(2, 5, 7, 7,1,2), 
  t5 = c(1, 0, 1, 1,1,1),
  t6 = c(1, 1, 1, 1,1,1), 
  t7 = c(1, 1, 1 ,1,1,1), 
  t8=c(0,0,0,0,0,0)
), row.names = c(NA, 6L), class = "data.frame")

1. Data Preparation: Focus on the t2 Metric

Your data has multiple rows per id (one for each se). First, we need to decide what to cluster:

  • Option 1: Cluster individual IDs (using their t2 values across all se levels)
  • Option 2: Cluster individual rows (each se entry for each id)

Let’s cover both, starting with the more common use case of clustering IDs.

Option 1: Reshape Data for ID-Level Clustering

We’ll pivot the data so each row is an ID, with columns representing t2 values for each se:

library(tidyr)
library(dplyr)

# Reshape to get one row per ID, with t2 values for each se
clust_data <- df %>%
  select(id, se, t2) %>%
  pivot_wider(names_from = se, values_from = t2, names_prefix = "t2_se_") %>%
  column_to_rownames("id") # Set ID as row names for easier clustering

# Check the result
clust_data

This gives us a clean dataset where each row is an ID, and columns are their t2 values across se=1,2,3.

Option 2: Use Individual Rows for Clustering

If you want to cluster the 6 rows directly using only t2, just extract the t2 column:

row_clust_data <- df %>% select(t2)

For the rest of this guide, we’ll use the ID-level clustered data (Option 1), but you can adapt the code to Option 2 easily.

2. Hierarchical Clustering with hclust()

Step 2.1: Calculate Distance Between Observations

First, compute a distance matrix (Euclidean distance is the default, but you can use Manhattan or others):

dist_matrix <- dist(clust_data, method = "euclidean")

Step 2.2: Run Hierarchical Clustering

Choose a linkage method—ward.D2 is great for minimizing within-cluster variance, but you can also try complete or single:

hier_clust <- hclust(dist_matrix, method = "ward.D2")

Step 2.3: Visualize the Dendrogram

Plot the dendrogram to see how observations group together:

plot(hier_clust, main = "Hierarchical Clustering Dendrogram (Based on t2)",
     xlab = "ID", ylab = "Distance", cex = 0.8)

Step 2.4: Find the Optimal Number of Clusters

Use these methods to pick the best k:

  • Elbow Method: Look for where the curve flattens (using the factoextra package):
library(factoextra)
fviz_nbclust(clust_data, FUN = hcut, method = "wss") +
  ggtitle("Elbow Method for Optimal Clusters")
  • Silhouette Method: Higher scores mean better cluster separation:
fviz_nbclust(clust_data, FUN = hcut, method = "silhouette") +
  ggtitle("Silhouette Method for Optimal Clusters")

Step 2.5: Assign Cluster Labels

Once you choose k (e.g., k=2), assign clusters to each ID:

k_optimal <- 2
hier_clusters <- cutree(hier_clust, k = k_optimal)

# Add cluster labels to your data
clust_data$hier_cluster <- hier_clusters

3. K-Means Clustering with kmeans()

Step 3.1: Run K-Means

Use the same optimal k from hierarchical clustering (or find it separately):

# Run k-means with 2 clusters, 20 initial starts to avoid local minima
kmeans_clust <- kmeans(clust_data %>% select(-hier_cluster), centers = k_optimal, nstart = 20)

Step 3.2: Analyze Results

Check key outputs from k-means:

# Cluster centers (average t2 values for each cluster)
kmeans_clust$centers

# Cluster assignments for each ID
kmeans_clust$cluster

# Total within-cluster variance (lower = better)
kmeans_clust$tot.withinss

Step 3.3: Assign K-Means Labels

Add the labels to your dataset:

clust_data$kmeans_cluster <- kmeans_clust$cluster

4. Compare Hierarchical & K-Means Results

See how the two methods assign clusters to IDs:

# View side-by-side cluster assignments
clust_data

# Create a contingency table to compare
table(clust_data$hier_cluster, clust_data$kmeans_cluster)

# Visualize the comparison
library(ggplot2)
clust_data %>%
  rownames_to_column("id") %>%
  ggplot(aes(x = hier_cluster, y = kmeans_cluster, color = id)) +
  geom_point(size = 4) +
  labs(title = "Hierarchical vs K-Means Cluster Assignments",
       x = "Hierarchical Cluster", y = "K-Means Cluster")

Quick Tips for Your Workflow

  • Data Shape: Always ensure each row is an observation (ID, row, etc.) and columns are features (like t2 values).
  • Method Selection: For hierarchical clustering, experiment with different linkage methods to see which fits your data best.
  • Optimal K: Don’t rely on one method—combine elbow/silhouette results with your domain knowledge to pick k.

内容的提问来源于stack exchange,提问作者Rstudent

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 23:37:52