You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

混合数据模糊聚类结果可视化技术问询(附R代码与数据样例)

Visualizing Fuzzy Clustering Results for Mixed Data

Hi there! Let's build on the work you've already done with Gower distance and t-SNE to create clear, informative visualizations for your fuzzy clustering results. Here's a step-by-step guide with code examples tailored to your mixed categorical + numerical data:

1. Fix & Complete Preprocessing Steps

I noticed a small typo in your Rtsne call (gower_dist1 should be gower_dist). Let's start by loading required packages and cleaning up the initial code for reproducibility:

# Load necessary packages
library(cluster)    # For daisy (Gower distance) and fanny (fuzzy clustering)
library(Rtsne)      # For t-SNE dimensionality reduction
library(ggplot2)    # For visualization
library(dplyr)      # For data manipulation
library(tidyr)      # For reshaping data

# Recreate your sample data
x <- data.frame(
  x1 = c("A", "B", "A", "B", "A", "B"),
  x2 = c("C", "C", "C", "D", "D", "C"),
  x3 = c(8.461373, 10.962334, 9.452127, 8.196687, 8.961367, 8.009029),
  x4 = c(27.62996, 27.22474, 27.57246, 27.29332, 26.72793, 27.97227),
  stringsAsFactors = FALSE
)

# Calculate Gower distance (ideal for mixed data types)
gower_dist <- daisy(x, metric = "gower")

# Run t-SNE to reduce to 2 dimensions
tsne_obj <- Rtsne(gower_dist, dims = 2, is_distance = TRUE)

2. Perform Fuzzy Clustering

Since you're working with fuzzy clustering, we'll use the fanny() function from the cluster package. Adjust the k value (number of clusters) based on your domain knowledge or validation metrics:

# Run fuzzy clustering (example with k=2 clusters)
fuzzy_clust <- fanny(gower_dist, k = 2, metric = "gower")

# Extract fuzzy membership values (each row = 1 sample, columns = membership to each cluster)
membership_df <- as.data.frame(fuzzy_clust$membership)
colnames(membership_df) <- paste0("Cluster_", 1:ncol(membership_df))

# Add a hard label for the cluster with the highest membership (for quick reference)
membership_df$dominant_cluster <- factor(apply(membership_df, 1, which.max))

3. Combine Data for Plotting

Merge the t-SNE coordinates, fuzzy memberships, and your original data into one dataframe for easy visualization:

# Create t-SNE coordinates dataframe
tsne_df <- data.frame(
  TSNE1 = tsne_obj$Y[,1],
  TSNE2 = tsne_obj$Y[,2]
)

# Combine all data into a single plot-ready dataframe
plot_data <- bind_cols(x, tsne_df, membership_df)

4. Visualization Options

Option 1: Scatter Plot with Fuzzy Membership Transparency

This plot uses color for the dominant cluster and transparency to show membership certainty (higher opacity = stronger membership to the dominant cluster):

ggplot(plot_data, aes(x = TSNE1, y = TSNE2)) +
  geom_point(
    aes(color = dominant_cluster, alpha = apply(select(plot_data, starts_with("Cluster_")), 1, max)),
    size = 3
  ) +
  scale_alpha(range = c(0.4, 1)) +
  labs(
    title = "t-SNE Visualization of Fuzzy Clustering Results",
    x = "t-SNE Dimension 1",
    y = "t-SNE Dimension 2",
    color = "Dominant Cluster",
    alpha = "Max Membership Value"
  ) +
  theme_minimal()

Option 2: Overlay Categorical Data

Highlight your original categorical variables (x1, x2) by adding shape or secondary color to the plot:

ggplot(plot_data, aes(x = TSNE1, y = TSNE2)) +
  geom_point(
    aes(color = dominant_cluster, shape = x1, alpha = apply(select(plot_data, starts_with("Cluster_")), 1, max)),
    size = 3
  ) +
  scale_alpha(range = c(0.4, 1)) +
  labs(
    title = "t-SNE with Fuzzy Clusters + Categorical Variable x1",
    x = "t-SNE Dimension 1",
    y = "t-SNE Dimension 2",
    color = "Dominant Cluster",
    shape = "x1 Category",
    alpha = "Max Membership Value"
  ) +
  theme_minimal()

Option 3: Fuzzy Membership Heatmap

Directly visualize membership distribution across clusters with a faceted heatmap:

# Reshape membership data for heatmap
membership_long <- plot_data %>%
  select(TSNE1, TSNE2, starts_with("Cluster_")) %>%
  pivot_longer(cols = starts_with("Cluster_"), names_to = "Cluster", values_to = "Membership")

ggplot(membership_long, aes(x = TSNE1, y = TSNE2)) +
  geom_tile(aes(fill = Membership)) +
  facet_wrap(~Cluster) +
  scale_fill_viridis_c(option = "plasma") +
  labs(
    title = "Fuzzy Membership Distribution Across Clusters",
    x = "t-SNE Dimension 1",
    y = "t-SNE Dimension 2",
    fill = "Membership Value"
  ) +
  theme_minimal()

Key Tips

  • Validate your cluster count k using metrics like the fuzzy silhouette coefficient (fuzzy_clust$silinfo$avg.width).
  • Gower distance is critical here because it handles both categorical and numerical data without requiring normalization.
  • Emphasize the fuzzy nature of your results (avoid treating them as hard clusters) – transparency and heatmaps are great ways to show membership uncertainty.

内容的提问来源于stack exchange,提问作者Jack shephard

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:31:15