You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中应用集成聚类?Scikit-learn相关实现咨询

Integrating Clustering Methods with Scikit-Learn for Your Dataset

Great question! Let's break this down clearly: Scikit-learn doesn't have dedicated built-in classes for ensemble clustering (unlike ensemble classification, where you have RandomForestClassifier or AdaBoostClassifier out of the box). But you can absolutely build powerful ensemble clustering pipelines using its existing tools. Here are the most common approaches with code examples tailored to your use case (since you already know K-Means, we'll use that as the base model):

1. Bagging-Based Ensemble Clustering (Easiest to Implement)

Bagging works by training multiple base clustering models on randomly resampled subsets of your data, then merging their results into a consensus. This reduces variance and improves robustness over a single K-Means run.

import numpy as np
from sklearn.cluster import KMeans
from sklearn.utils import resample
from sklearn.metrics import adjusted_rand_score
from sklearn.cluster import AgglomerativeClustering

# Replace this with your actual dataset (shape: [n_samples, n_features])
X = ...
# Define your parameters
n_estimators = 10  # Number of K-Means models to train
n_target_clusters = 5  # Your desired number of clusters
sample_fraction = 0.8  # Fraction of data to sample each time (with replacement)

# Store labels from each base model
cluster_label_list = []

# Train multiple K-Means models on resampled data
for _ in range(n_estimators):
    # Resample the dataset
    X_resampled = resample(X, replace=True, n_samples=int(len(X)*sample_fraction))
    # Train K-Means on the resampled data
    kmeans = KMeans(n_clusters=n_target_clusters, random_state=42)
    kmeans.fit(X_resampled)
    # Predict labels for the original dataset
    labels = kmeans.predict(X)
    cluster_label_list.append(labels)

# Merge results: treat each model's labels as a new feature, then cluster those features
label_matrix = np.array(cluster_label_list).T
# Use hierarchical clustering with Hamming distance (for discrete label features)
consensus_clusterer = AgglomerativeClustering(
    n_clusters=n_target_clusters,
    metric='hamming',
    linkage='average'
)
final_ensemble_labels = consensus_clusterer.fit_predict(label_matrix)

# Compare performance to a single K-Means run (if you have true labels)
single_kmeans = KMeans(n_clusters=n_target_clusters, random_state=42)
single_run_labels = single_kmeans.fit_predict(X)
if 'true_labels' in locals():
    print(f"Single K-Means ARI: {adjusted_rand_score(true_labels, single_run_labels):.3f}")
    print(f"Bagged Ensemble ARI: {adjusted_rand_score(true_labels, final_ensemble_labels):.3f}")

2. Boosting-Based Ensemble Clustering

Boosting focuses on iteratively improving the model by weighting samples that were poorly clustered in previous runs. For clustering, we use the distance from each sample to its cluster center as a proxy for "uncertainty" (higher distance = more uncertain, so we weight these samples more in the next iteration).

import numpy as np
from sklearn.cluster import KMeans
from sklearn.utils import resample

X = ...
n_estimators = 5
n_target_clusters = 5
# Initialize uniform sample weights
sample_weights = np.ones(len(X)) / len(X)

# Accumulate weighted labels
weighted_label_sum = np.zeros(len(X))

for _ in range(n_estimators):
    # Resample data based on current weights
    X_resampled, indices = resample(
        X,
        replace=True,
        n_samples=len(X),
        weights=sample_weights,
        return_indices=True
    )
    # Train K-Means
    kmeans = KMeans(n_clusters=n_target_clusters, random_state=42)
    kmeans.fit(X_resampled)
    labels = kmeans.predict(X)
    # Calculate uncertainty: distance to the nearest cluster center
    sample_distances = kmeans.transform(X).min(axis=1)
    # Update weights: higher distance = higher weight
    sample_weights = sample_distances / sample_distances.sum()
    # Accumulate labels with weights
    weighted_label_sum += labels * sample_weights

# Convert weighted sums to final cluster labels
final_boosted_labels = np.round(weighted_label_sum).astype(int)

3. Consensus Clustering (Co-Association Matrix Approach)

This method builds a co-association matrix that counts how often each pair of samples is assigned to the same cluster across all base models. We then cluster this similarity matrix to get the final result.

import numpy as np
from sklearn.cluster import KMeans
from sklearn.utils import resample
from sklearn.cluster import AgglomerativeClustering

X = ...
n_estimators = 10
n_target_clusters = 5

# Initialize co-association matrix (counts co-clustering occurrences)
co_assoc_matrix = np.zeros((len(X), len(X)))

for _ in range(n_estimators):
    X_resampled = resample(X, replace=True)
    kmeans = KMeans(n_clusters=n_target_clusters, random_state=42)
    labels = kmeans.predict(X)
    # Update the matrix: increment count for pairs in the same cluster
    for i in range(len(X)):
        for j in range(i + 1, len(X)):
            if labels[i] == labels[j]:
                co_assoc_matrix[i, j] += 1
                co_assoc_matrix[j, i] += 1

# Normalize to get a similarity score (0 to 1)
co_assoc_matrix /= n_estimators
# Convert to a distance matrix (since clustering uses distances)
distance_matrix = 1 - co_assoc_matrix

# Cluster the distance matrix with hierarchical clustering
consensus_clusterer = AgglomerativeClustering(
    n_clusters=n_target_clusters,
    metric='precomputed',
    linkage='average'
)
final_consensus_labels = consensus_clusterer.fit_predict(distance_matrix)

Quick Tips

  • Evaluation: If you have ground-truth labels, use metrics like adjusted_rand_score or normalized_mutual_info_score to compare ensemble performance against single models.
  • Flexibility: Replace K-Means with other base clusterers like DBSCAN or Gaussian Mixture Models—just make sure they can output cluster labels for all samples.
  • Alternative Libraries: If you don't want to build from scratch, libraries like cluster-ensemble exist, but scikit-learn's tools are more than sufficient for most use cases.

内容的提问来源于stack exchange,提问作者IS2057

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 03:29:47