如何生成预定义簇的树状图并对比其与层次聚类结果?
Hey there! Let's work through how to compare your predefined clusters with the agnes hierarchical clustering results—including creating a dendrogram for your manual groups and other handy comparison methods. First, let's fix a small issue in your original code: your c1 and c2 columns are stored as characters, so we'll convert them to numeric first for valid distance calculations.
Required Packages
Make sure you have these packages loaded to run all the code below:
library(dplyr) library(factoextra) library(clValid) library(dendextend) library(aricode) library(ggplot2) library(FactoMineR)
Fix Your Data Type
set.seed(73) great <- data.frame( c0=c("r1","r2","r3","r4","r5","r6"), c1=c("0.89","46","0","0.56","12","0"), c2=c("0","0.45","45","79","0.45","4.4") ) %>% mutate(across(c(c1, c2), as.numeric)) # Convert to numeric for distance calculations
1. Create a Dendrogram for Predefined Clusters
To use tanglegram(), we need an hclust object for your manual clusters. Since your groups are already fixed (cl1: r1/r2/r4; cl2: r3/r5/r6), we'll manually build a hierarchy that reflects this grouping.
# Define predefined cluster labels as a numeric vector predefined_clusters <- rep(2, nrow(great)) # Default to cl2 predefined_clusters[c(1,2,4)] <- 1 # Set cl1 for r1, r2, r4 # Build a manual hclust object pre_hclust <- list( # Merge matrix: negatives = individual samples, positives = merged clusters merge = matrix( c(-1, -2, # Merge r1 & r2 first -4, 1, # Merge r4 with the r1-r2 cluster (now cluster 1) -3, -5, # Merge r3 & r5 first -6, 3, # Merge r6 with the r3-r5 cluster (now cluster 3) 2, 4), # Merge the two final clusters (cluster 2 & 4) ncol = 2, byrow = TRUE ), # Heights: arbitrary values to represent hierarchy levels height = c(0.1, 0.2, 0.1, 0.2, 1), # Order of samples in the dendrogram order = c(1,2,4,3,5,6), # Sample labels labels = great$c0 ) class(pre_hclust) <- "hclust" # Assign hclust class to the object
Generate the Tanglegram
Now compare your agnes result with the predefined clusters using tanglegram():
# Convert your agnes result to hclust format agnes_hclust <- as.hclust(agnes(great, method = "ward")) # Create the tanglegram tanglegram( agnes_hclust, pre_hclust, main_left = "Agnes (Ward's Method)", main_right = "Predefined Clusters", cex = 0.7, common_subtrees_color_lines = TRUE, highlight_distinct_edges = TRUE )
This plot will align the two dendrograms and visually highlight where sample groupings differ or match between the two clusterings.
2. Other Cluster Comparison Methods
Beyond tanglegrams, here are practical ways to quantify and visualize differences between your clusterings:
Adjusted Rand Index (ARI)
A normalized measure of similarity between two clusterings, ranging from -1 (complete disagreement) to 1 (perfect agreement). It accounts for chance, making it more reliable than raw counts.
ari_score <- ARI(predefined_clusters, cutree(agnes_hclust, k=2)) cat("Adjusted Rand Index:", round(ari_score, 3), "\n")
Contingency Table
A simple way to see how samples map between the two clusterings:
cluster_table <- table( Predefined = factor(predefined_clusters, labels = c("cl1", "cl2")), Agnes = factor(cutree(agnes_hclust, k=2), labels = c("Cluster 1", "Cluster 2")) ) print(cluster_table)
Silhouette Analysis
Compare how well each clustering separates samples. Higher average silhouette scores indicate better cluster structure:
great_dist <- dist(great[, -1]) # Silhouette for agnes clusters sil_agnes <- silhouette(cutree(agnes_hclust, k=2), great_dist) fviz_silhouette(sil_agnes, main = "Silhouette Plot: Agnes Ward Clustering") # Silhouette for predefined clusters sil_pre <- silhouette(predefined_clusters, great_dist) fviz_silhouette(sil_pre, main = "Silhouette Plot: Predefined Clusters")
PCA Visualization
Plot your data in 2D using PCA, then color points by both clusterings to visually compare groupings:
# Run PCA (exclude the ID column) pca_res <- PCA(great[, -1], graph = FALSE) pca_data <- cbind(great[, 1], as.data.frame(pca_res$ind$coord)) colnames(pca_data)[1] <- "ID" pca_data$Predefined <- factor(predefined_clusters, labels = c("cl1", "cl2")) pca_data$Agnes <- factor(cutree(agnes_hclust, k=2), labels = c("Cluster 1", "Cluster 2")) # Plot with predefined clusters ggplot(pca_data, aes(PC1, PC2, color = Predefined, label = ID)) + geom_point(size = 4) + geom_text(nudge_y = 0.1) + labs(title = "PCA Plot: Predefined Clusters") + theme_minimal() # Plot with agnes clusters ggplot(pca_data, aes(PC1, PC2, color = Agnes, label = ID)) + geom_point(size = 4) + geom_text(nudge_y = 0.1) + labs(title = "PCA Plot: Agnes Ward Clustering") + theme_minimal()
内容的提问来源于stack exchange,提问作者takeITeasy

