You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何阻止R语言clustMixType包kproto函数合并相同原型?

Fixing k-prototype Cluster Merging for labelMatcher() Compatibility

Great question—this is a common gotcha with the kproto() function from clustMixType! The issue happens because by default, kproto() will automatically drop clusters that end up empty (or with identical prototypes) during its iterative process, leaving you with fewer unique cluster labels than you specified. This breaks labelMatcher() since it expects the number of clusters to match your number of true classes.

Here are two straightforward fixes to prevent this merging:

1. Force Keep Empty Clusters with keep.empty = TRUE

The simplest solution is to use the keep.empty parameter in kproto(), which tells the function to retain all k clusters you specify—even if some end up with no assigned samples during iteration. This ensures you get exactly k unique cluster labels, which labelMatcher() can work with.

Example code:

# Load required packages
library(clustMixType)
library(Thresher)

# Define your target number of clusters (match your true class count)
target_k <- 3 # Adjust this to your hepatitis dataset's class count

# Run k-prototype clustering with keep.empty enabled
clust_output <- kproto(
  data = your_hepatitis_dataset,
  k = target_k,
  keep.empty = TRUE,
  init = "random" # Random initialization helps avoid identical starting prototypes
)

# Verify you have the correct number of unique labels
length(unique(clust_output$cluster)) # Should equal target_k

2. Manually Specify Initial Prototypes

If you still run into identical prototypes (even with keep.empty = TRUE), the problem might be with random initialization creating duplicate starting points. In this case, manually select distinct initial prototypes from your dataset to ensure all k clusters start as unique.

Example code:

# Pick k distinct rows from your dataset as initial prototypes
set.seed(123) # For reproducibility
initial_prototypes <- your_hepatitis_dataset[sample(nrow(your_hepatitis_dataset), target_k), ]

# Run k-prototype with manual initialization
clust_output <- kproto(
  data = your_hepatitis_dataset,
  k = target_k,
  keep.empty = TRUE,
  init = initial_prototypes
)

Post-Clustering Cleanup (If Needed)

If you end up with empty clusters after using keep.empty = TRUE, you can either:

  • Re-run the clustering with a different random seed (to avoid empty clusters), or
  • Assign those empty cluster labels to samples manually (though re-running with better initialization is the more reliable approach).

Once you have the correct number of cluster labels, you can proceed with labelMatcher() as usual:

# Match cluster labels to true class labels
match_result <- labelMatcher(
  predicted = clust_output$cluster,
  actual = your_true_hepatitis_labels
)

# Extract evaluation metrics
accuracy <- match_result$accuracy
sensitivity <- match_result$sensitivity
precision <- match_result$precision
f1_score <- match_result$f1

内容的提问来源于stack exchange,提问作者Mehmet Yildirim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:31:48