混合数据模糊聚类结果可视化技术问询(附R代码与数据样例)
Hi there! Let's build on the work you've already done with Gower distance and t-SNE to create clear, informative visualizations for your fuzzy clustering results. Here's a step-by-step guide with code examples tailored to your mixed categorical + numerical data:
1. Fix & Complete Preprocessing Steps
I noticed a small typo in your Rtsne call (gower_dist1 should be gower_dist). Let's start by loading required packages and cleaning up the initial code for reproducibility:
# Load necessary packages library(cluster) # For daisy (Gower distance) and fanny (fuzzy clustering) library(Rtsne) # For t-SNE dimensionality reduction library(ggplot2) # For visualization library(dplyr) # For data manipulation library(tidyr) # For reshaping data # Recreate your sample data x <- data.frame( x1 = c("A", "B", "A", "B", "A", "B"), x2 = c("C", "C", "C", "D", "D", "C"), x3 = c(8.461373, 10.962334, 9.452127, 8.196687, 8.961367, 8.009029), x4 = c(27.62996, 27.22474, 27.57246, 27.29332, 26.72793, 27.97227), stringsAsFactors = FALSE ) # Calculate Gower distance (ideal for mixed data types) gower_dist <- daisy(x, metric = "gower") # Run t-SNE to reduce to 2 dimensions tsne_obj <- Rtsne(gower_dist, dims = 2, is_distance = TRUE)
2. Perform Fuzzy Clustering
Since you're working with fuzzy clustering, we'll use the fanny() function from the cluster package. Adjust the k value (number of clusters) based on your domain knowledge or validation metrics:
# Run fuzzy clustering (example with k=2 clusters) fuzzy_clust <- fanny(gower_dist, k = 2, metric = "gower") # Extract fuzzy membership values (each row = 1 sample, columns = membership to each cluster) membership_df <- as.data.frame(fuzzy_clust$membership) colnames(membership_df) <- paste0("Cluster_", 1:ncol(membership_df)) # Add a hard label for the cluster with the highest membership (for quick reference) membership_df$dominant_cluster <- factor(apply(membership_df, 1, which.max))
3. Combine Data for Plotting
Merge the t-SNE coordinates, fuzzy memberships, and your original data into one dataframe for easy visualization:
# Create t-SNE coordinates dataframe tsne_df <- data.frame( TSNE1 = tsne_obj$Y[,1], TSNE2 = tsne_obj$Y[,2] ) # Combine all data into a single plot-ready dataframe plot_data <- bind_cols(x, tsne_df, membership_df)
4. Visualization Options
Option 1: Scatter Plot with Fuzzy Membership Transparency
This plot uses color for the dominant cluster and transparency to show membership certainty (higher opacity = stronger membership to the dominant cluster):
ggplot(plot_data, aes(x = TSNE1, y = TSNE2)) + geom_point( aes(color = dominant_cluster, alpha = apply(select(plot_data, starts_with("Cluster_")), 1, max)), size = 3 ) + scale_alpha(range = c(0.4, 1)) + labs( title = "t-SNE Visualization of Fuzzy Clustering Results", x = "t-SNE Dimension 1", y = "t-SNE Dimension 2", color = "Dominant Cluster", alpha = "Max Membership Value" ) + theme_minimal()
Option 2: Overlay Categorical Data
Highlight your original categorical variables (x1, x2) by adding shape or secondary color to the plot:
ggplot(plot_data, aes(x = TSNE1, y = TSNE2)) + geom_point( aes(color = dominant_cluster, shape = x1, alpha = apply(select(plot_data, starts_with("Cluster_")), 1, max)), size = 3 ) + scale_alpha(range = c(0.4, 1)) + labs( title = "t-SNE with Fuzzy Clusters + Categorical Variable x1", x = "t-SNE Dimension 1", y = "t-SNE Dimension 2", color = "Dominant Cluster", shape = "x1 Category", alpha = "Max Membership Value" ) + theme_minimal()
Option 3: Fuzzy Membership Heatmap
Directly visualize membership distribution across clusters with a faceted heatmap:
# Reshape membership data for heatmap membership_long <- plot_data %>% select(TSNE1, TSNE2, starts_with("Cluster_")) %>% pivot_longer(cols = starts_with("Cluster_"), names_to = "Cluster", values_to = "Membership") ggplot(membership_long, aes(x = TSNE1, y = TSNE2)) + geom_tile(aes(fill = Membership)) + facet_wrap(~Cluster) + scale_fill_viridis_c(option = "plasma") + labs( title = "Fuzzy Membership Distribution Across Clusters", x = "t-SNE Dimension 1", y = "t-SNE Dimension 2", fill = "Membership Value" ) + theme_minimal()
Key Tips
- Validate your cluster count
kusing metrics like the fuzzy silhouette coefficient (fuzzy_clust$silinfo$avg.width). - Gower distance is critical here because it handles both categorical and numerical data without requiring normalization.
- Emphasize the fuzzy nature of your results (avoid treating them as hard clusters) – transparency and heatmaps are great ways to show membership uncertainty.
内容的提问来源于stack exchange,提问作者Jack shephard

