You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

CLARA无Gower距离选项的原因及解决方案(混合类型大数据聚类)

Great question! Let's break this down into two parts: why CLARA implementations in cluster and ClusterR don't support Gower distance, and how you can work around this for your large mixed-type dataset.

Why no Gower option in standard CLARA functions?

  • Historical algorithm design: The original CLARA algorithm was built as a scalable extension of PAM (Partitioning Around Medoids), which was initially designed for numerical data using Euclidean or Manhattan distances. The cluster package's implementation sticks close to this original design, so it didn't include support for mixed-type distances like Gower.
  • Performance tradeoffs: Gower distance requires custom handling for different variable types (e.g., matching nominal variables, scaling continuous ones, ranking ordered variables). Adding this to CLARA's sampling and medoid computation pipeline would introduce extra complexity and runtime overhead, which the maintainers of cluster and ClusterR likely prioritized avoiding for their target use cases (numerical data at scale).

How to cluster your large mixed-type dataset with CLARA-like logic

You have two solid options here: using a dedicated mixed-type clustering package, or rolling your own simplified CLARA workflow with Gower distance.

Option 1: Use clustMixType (simplest, most robust)

The clustMixType package was specifically built for clustering mixed-type data, and it includes a CLARA implementation that automatically handles continuous, ordered, and nominal variables using a Gower-like distance metric. No need to manually compute distances—it all happens under the hood.

Here's how to use it:

# Install and load the package
install.packages("clustMixType")
library(clustMixType)

# Run CLARA on your dataset (replace k with your desired number of clusters)
# samples = number of random samples to draw; sampsize = size of each sample
clara_mixed <- clara(your_dataframe, k = 5, samples = 50, sampsize = 1000)

# Get cluster labels for all 11.4M records
cluster_labels <- clara_mixed$cluster

This implementation follows the core CLARA logic (sampling, clustering samples, assigning rest of data to medoids) but is optimized for mixed-type data. It's by far the easiest way to get the result you want.

Option 2: Roll your own CLARA workflow with Gower distance

If you prefer to stick with base cluster and gower packages, you can replicate CLARA's logic manually. The idea is to:

  1. Draw multiple random samples from your large dataset
  2. Compute Gower distances for each sample, run PAM to get medoids
  3. Assign all data points to the closest medoid from each sample
  4. Keep the clustering result with the lowest total dissimilarity

Here's a code example:

library(gower)
library(cluster)
set.seed(123) # For reproducibility

# Define parameters
num_clusters <- 5
num_samples <- 50 # Number of CLARA samples to test
sample_size <- 1000 # Size of each sample

best_total_diss <- Inf
best_clusters <- NULL

for (i in 1:num_samples) {
  # Step 1: Draw a random sample
  sample_indices <- sample(nrow(your_dataframe), sample_size)
  sample_data <- your_dataframe[sample_indices, ]
  
  # Step 2: Compute Gower distance for the sample and run PAM
  sample_gower <- gower_dist(sample_data)
  pam_fit <- pam(sample_gower, k = num_clusters, diss = TRUE)
  
  # Step 3: Get medoids from the sample
  medoids <- sample_data[pam_fit$medoids, ]
  
  # Step 4: Compute distance from all data to medoids, assign clusters
  all_distances <- gower_dist(your_dataframe, medoids)
  current_clusters <- apply(all_distances, 1, which.min)
  
  # Step 5: Calculate total dissimilarity, keep the best result
  total_diss <- sum(apply(all_distances, 1, min))
  if (total_diss < best_total_diss) {
    best_total_diss <- total_diss
    best_clusters <- current_clusters
  }
}

# best_clusters now holds the optimal cluster labels for all records

This gives you full control over the sampling and clustering process, though it will take longer to run than the clustMixType approach.

内容的提问来源于stack exchange,提问作者Menn

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 07:34:16