You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言apcluster的聚类可视化模板实现需求

Using apcluster in R for Template-Based Visualization of Clustering Results

Hey there! Let's walk through how to use the apcluster package in R to perform affinity propagation clustering and create reusable visualization templates for your results. I'll use your provided data structure and reading code as a starting point.


Step 1: Install & Load Required Packages

First, make sure you have apcluster installed. If not, grab it from CRAN:

# Install package if needed
if (!require("apcluster")) {
  install.packages("apcluster")
  library(apcluster)
}
# Optional: Use ggplot2 for enhanced visuals (if needed)
if (!require("ggplot2")) {
  install.packages("ggplot2")
  library(ggplot2)
}

Step 2: Load & Inspect Your Data

Use your provided code to load the CSV, and let's confirm the data structure with the sample you shared:

# Load data (adjust path/sep/dec as needed)
mydat <- read.csv("mydat.csv", sep = ";", dec = ",")

# If you want to test with your sample structure, use this:
mydat <- structure(
  list(
    x1 = c(0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 0L, 0L),
    x2 = c(0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 1L, 1L, 1L, 1L, 1L, 0L, 0L, 0L),
    x3 = c(0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L),
    x4 = c(0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L, 1L, 0L, 0L, 0L, 0L, 0L)
    # Add remaining columns from your dput snippet here if needed
  ),
  class = "data.frame",
  row.names = c(NA, -21L)
)

# Check data structure to confirm it loaded correctly
str(mydat)

Step 3: Run Affinity Propagation Clustering

Next, we'll compute the affinity matrix and run the clustering. For binary data like yours, the Jaccard similarity is a great choice:

# Compute Jaccard similarity matrix (ideal for binary feature data)
affinity_mat <- negDistMat(mydat, method = "jaccard")

# Run affinity propagation clustering
ap_result <- apcluster(affinity_mat, q = 0.1) # Adjust q to control number of clusters

Note: The q parameter sets the preference for how many clusters are formed—higher values lead to more clusters. Play around with it to get the right number for your dataset.

Step 4: Template-Based Visualization

Let's create reusable visualization functions (templates) so you can quickly plot results any time you run the clustering.

Template 1: Cluster Membership Bar Plot

This template shows which cluster each observation belongs to, sorted for easy interpretation:

# Define reusable plot function
plot_cluster_membership <- function(ap_result, data) {
  # Create dataframe with cluster labels
  cluster_df <- data.frame(
    Observation = rownames(data),
    Cluster = factor(ap_result@clusters)
  )
  
  # Generate plot
  ggplot(cluster_df, aes(x = reorder(Observation, Cluster), y = 1, fill = Cluster)) +
    geom_col(width = 1) +
    labs(title = "Cluster Membership of Observations", x = "Observation", y = "") +
    theme_minimal() +
    theme(axis.text.y = element_blank(), axis.ticks.y = element_blank())
}

# Use the template with your results
plot_cluster_membership(ap_result, mydat)

Template 2: Clustered Data Heatmap

This template visualizes your original data grouped by clusters, highlighting patterns across features:

# Define reusable plot function
plot_cluster_heatmap <- function(ap_result, data) {
  # Reorder data by cluster membership
  ordered_data <- data[order(ap_result@clusters), ]
  cluster_labels <- ap_result@clusters[order(ap_result@clusters)]
  
  # Generate heatmap with cluster side colors
  heatmap(
    as.matrix(ordered_data),
    Colv = NA,
    Rowv = NA,
    col = c("white", "black"),
    main = "Heatmap of Data Grouped by Clusters",
    RowSideColors = rainbow(length(unique(cluster_labels)))[cluster_labels]
  )
  # Add legend for clusters
  legend(
    "topright",
    legend = unique(cluster_labels),
    fill = rainbow(length(unique(cluster_labels))),
    title = "Clusters"
  )
}

# Use the template with your results
plot_cluster_heatmap(ap_result, mydat)

Template 3: PCA Plot with Exemplar Highlighting

Affinity propagation identifies exemplars (representative points) for each cluster. This template uses PCA to reduce dimensionality and highlights these key points:

# Define reusable plot function
plot_exemplars <- function(ap_result, data) {
  # Get exemplar indices and labels
  exemplar_indices <- ap_result@exemplars
  exemplar_labels <- rownames(data)[exemplar_indices]
  
  # Create dataframe with cluster and exemplar status
  plot_df <- data.frame(
    data,
    Cluster = factor(ap_result@clusters),
    IsExemplar = FALSE
  )
  plot_df$IsExemplar[exemplar_indices] <- TRUE
  
  # Reduce dimensions with PCA for visualization
  pca <- prcomp(data, scale. = TRUE)
  plot_df$PC1 <- pca$x[, 1]
  plot_df$PC2 <- pca$x[, 2]
  
  # Generate plot
  ggplot(plot_df, aes(x = PC1, y = PC2, color = Cluster, shape = IsExemplar)) +
    geom_point(size = 3) +
    labs(title = "PCA Plot with Cluster Exemplars", x = "Principal Component 1", y = "Principal Component 2") +
    theme_minimal() +
    scale_shape_manual(values = c(16, 17)) # Circle = regular point, Triangle = exemplar
}

# Use the template with your results
plot_exemplars(ap_result, mydat)

Quick Customization Tips

  • Adjust the q parameter in apcluster() to tweak the number of clusters.
  • Swap out color palettes (e.g., use scale_fill_viridis_d() instead of rainbow for better accessibility).
  • For binary data, try Hamming distance in negDistMat() as an alternative similarity metric.

内容的提问来源于stack exchange,提问作者psysky

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:38:44