基于R语言apcluster的聚类可视化模板实现需求
Hey there! Let's walk through how to use the apcluster package in R to perform affinity propagation clustering and create reusable visualization templates for your results. I'll use your provided data structure and reading code as a starting point.
Step 1: Install & Load Required Packages
First, make sure you have apcluster installed. If not, grab it from CRAN:
# Install package if needed if (!require("apcluster")) { install.packages("apcluster") library(apcluster) } # Optional: Use ggplot2 for enhanced visuals (if needed) if (!require("ggplot2")) { install.packages("ggplot2") library(ggplot2) }
Step 2: Load & Inspect Your Data
Use your provided code to load the CSV, and let's confirm the data structure with the sample you shared:
# Load data (adjust path/sep/dec as needed) mydat <- read.csv("mydat.csv", sep = ";", dec = ",") # If you want to test with your sample structure, use this: mydat <- structure( list( x1 = c(0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 0L, 0L), x2 = c(0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 0L, 0L), x3 = c(0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L), x4 = c(0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L) # Add remaining columns from your dput snippet here if needed ), class = "data.frame", row.names = c(NA, -21L) ) # Check data structure to confirm it loaded correctly str(mydat)
Step 3: Run Affinity Propagation Clustering
Next, we'll compute the affinity matrix and run the clustering. For binary data like yours, the Jaccard similarity is a great choice:
# Compute Jaccard similarity matrix (ideal for binary feature data) affinity_mat <- negDistMat(mydat, method = "jaccard") # Run affinity propagation clustering ap_result <- apcluster(affinity_mat, q = 0.1) # Adjust q to control number of clusters
Note: The q parameter sets the preference for how many clusters are formed—higher values lead to more clusters. Play around with it to get the right number for your dataset.
Step 4: Template-Based Visualization
Let's create reusable visualization functions (templates) so you can quickly plot results any time you run the clustering.
Template 1: Cluster Membership Bar Plot
This template shows which cluster each observation belongs to, sorted for easy interpretation:
# Define reusable plot function plot_cluster_membership <- function(ap_result, data) { # Create dataframe with cluster labels cluster_df <- data.frame( Observation = rownames(data), Cluster = factor(ap_result@clusters) ) # Generate plot ggplot(cluster_df, aes(x = reorder(Observation, Cluster), y = 1, fill = Cluster)) + geom_col(width = 1) + labs(title = "Cluster Membership of Observations", x = "Observation", y = "") + theme_minimal() + theme(axis.text.y = element_blank(), axis.ticks.y = element_blank()) } # Use the template with your results plot_cluster_membership(ap_result, mydat)
Template 2: Clustered Data Heatmap
This template visualizes your original data grouped by clusters, highlighting patterns across features:
# Define reusable plot function plot_cluster_heatmap <- function(ap_result, data) { # Reorder data by cluster membership ordered_data <- data[order(ap_result@clusters), ] cluster_labels <- ap_result@clusters[order(ap_result@clusters)] # Generate heatmap with cluster side colors heatmap( as.matrix(ordered_data), Colv = NA, Rowv = NA, col = c("white", "black"), main = "Heatmap of Data Grouped by Clusters", RowSideColors = rainbow(length(unique(cluster_labels)))[cluster_labels] ) # Add legend for clusters legend( "topright", legend = unique(cluster_labels), fill = rainbow(length(unique(cluster_labels))), title = "Clusters" ) } # Use the template with your results plot_cluster_heatmap(ap_result, mydat)
Template 3: PCA Plot with Exemplar Highlighting
Affinity propagation identifies exemplars (representative points) for each cluster. This template uses PCA to reduce dimensionality and highlights these key points:
# Define reusable plot function plot_exemplars <- function(ap_result, data) { # Get exemplar indices and labels exemplar_indices <- ap_result@exemplars exemplar_labels <- rownames(data)[exemplar_indices] # Create dataframe with cluster and exemplar status plot_df <- data.frame( data, Cluster = factor(ap_result@clusters), IsExemplar = FALSE ) plot_df$IsExemplar[exemplar_indices] <- TRUE # Reduce dimensions with PCA for visualization pca <- prcomp(data, scale. = TRUE) plot_df$PC1 <- pca$x[, 1] plot_df$PC2 <- pca$x[, 2] # Generate plot ggplot(plot_df, aes(x = PC1, y = PC2, color = Cluster, shape = IsExemplar)) + geom_point(size = 3) + labs(title = "PCA Plot with Cluster Exemplars", x = "Principal Component 1", y = "Principal Component 2") + theme_minimal() + scale_shape_manual(values = c(16, 17)) # Circle = regular point, Triangle = exemplar } # Use the template with your results plot_exemplars(ap_result, mydat)
Quick Customization Tips
- Adjust the
qparameter inapcluster()to tweak the number of clusters. - Swap out color palettes (e.g., use
scale_fill_viridis_d()instead of rainbow for better accessibility). - For binary data, try Hamming distance in
negDistMat()as an alternative similarity metric.
内容的提问来源于stack exchange,提问作者psysky

