如何在Python中应用集成聚类?Scikit-learn相关实现咨询
Great question! Let's break this down clearly: Scikit-learn doesn't have dedicated built-in classes for ensemble clustering (unlike ensemble classification, where you have RandomForestClassifier or AdaBoostClassifier out of the box). But you can absolutely build powerful ensemble clustering pipelines using its existing tools. Here are the most common approaches with code examples tailored to your use case (since you already know K-Means, we'll use that as the base model):
1. Bagging-Based Ensemble Clustering (Easiest to Implement)
Bagging works by training multiple base clustering models on randomly resampled subsets of your data, then merging their results into a consensus. This reduces variance and improves robustness over a single K-Means run.
import numpy as np from sklearn.cluster import KMeans from sklearn.utils import resample from sklearn.metrics import adjusted_rand_score from sklearn.cluster import AgglomerativeClustering # Replace this with your actual dataset (shape: [n_samples, n_features]) X = ... # Define your parameters n_estimators = 10 # Number of K-Means models to train n_target_clusters = 5 # Your desired number of clusters sample_fraction = 0.8 # Fraction of data to sample each time (with replacement) # Store labels from each base model cluster_label_list = [] # Train multiple K-Means models on resampled data for _ in range(n_estimators): # Resample the dataset X_resampled = resample(X, replace=True, n_samples=int(len(X)*sample_fraction)) # Train K-Means on the resampled data kmeans = KMeans(n_clusters=n_target_clusters, random_state=42) kmeans.fit(X_resampled) # Predict labels for the original dataset labels = kmeans.predict(X) cluster_label_list.append(labels) # Merge results: treat each model's labels as a new feature, then cluster those features label_matrix = np.array(cluster_label_list).T # Use hierarchical clustering with Hamming distance (for discrete label features) consensus_clusterer = AgglomerativeClustering( n_clusters=n_target_clusters, metric='hamming', linkage='average' ) final_ensemble_labels = consensus_clusterer.fit_predict(label_matrix) # Compare performance to a single K-Means run (if you have true labels) single_kmeans = KMeans(n_clusters=n_target_clusters, random_state=42) single_run_labels = single_kmeans.fit_predict(X) if 'true_labels' in locals(): print(f"Single K-Means ARI: {adjusted_rand_score(true_labels, single_run_labels):.3f}") print(f"Bagged Ensemble ARI: {adjusted_rand_score(true_labels, final_ensemble_labels):.3f}")
2. Boosting-Based Ensemble Clustering
Boosting focuses on iteratively improving the model by weighting samples that were poorly clustered in previous runs. For clustering, we use the distance from each sample to its cluster center as a proxy for "uncertainty" (higher distance = more uncertain, so we weight these samples more in the next iteration).
import numpy as np from sklearn.cluster import KMeans from sklearn.utils import resample X = ... n_estimators = 5 n_target_clusters = 5 # Initialize uniform sample weights sample_weights = np.ones(len(X)) / len(X) # Accumulate weighted labels weighted_label_sum = np.zeros(len(X)) for _ in range(n_estimators): # Resample data based on current weights X_resampled, indices = resample( X, replace=True, n_samples=len(X), weights=sample_weights, return_indices=True ) # Train K-Means kmeans = KMeans(n_clusters=n_target_clusters, random_state=42) kmeans.fit(X_resampled) labels = kmeans.predict(X) # Calculate uncertainty: distance to the nearest cluster center sample_distances = kmeans.transform(X).min(axis=1) # Update weights: higher distance = higher weight sample_weights = sample_distances / sample_distances.sum() # Accumulate labels with weights weighted_label_sum += labels * sample_weights # Convert weighted sums to final cluster labels final_boosted_labels = np.round(weighted_label_sum).astype(int)
3. Consensus Clustering (Co-Association Matrix Approach)
This method builds a co-association matrix that counts how often each pair of samples is assigned to the same cluster across all base models. We then cluster this similarity matrix to get the final result.
import numpy as np from sklearn.cluster import KMeans from sklearn.utils import resample from sklearn.cluster import AgglomerativeClustering X = ... n_estimators = 10 n_target_clusters = 5 # Initialize co-association matrix (counts co-clustering occurrences) co_assoc_matrix = np.zeros((len(X), len(X))) for _ in range(n_estimators): X_resampled = resample(X, replace=True) kmeans = KMeans(n_clusters=n_target_clusters, random_state=42) labels = kmeans.predict(X) # Update the matrix: increment count for pairs in the same cluster for i in range(len(X)): for j in range(i + 1, len(X)): if labels[i] == labels[j]: co_assoc_matrix[i, j] += 1 co_assoc_matrix[j, i] += 1 # Normalize to get a similarity score (0 to 1) co_assoc_matrix /= n_estimators # Convert to a distance matrix (since clustering uses distances) distance_matrix = 1 - co_assoc_matrix # Cluster the distance matrix with hierarchical clustering consensus_clusterer = AgglomerativeClustering( n_clusters=n_target_clusters, metric='precomputed', linkage='average' ) final_consensus_labels = consensus_clusterer.fit_predict(distance_matrix)
Quick Tips
- Evaluation: If you have ground-truth labels, use metrics like
adjusted_rand_scoreornormalized_mutual_info_scoreto compare ensemble performance against single models. - Flexibility: Replace K-Means with other base clusterers like DBSCAN or Gaussian Mixture Models—just make sure they can output cluster labels for all samples.
- Alternative Libraries: If you don't want to build from scratch, libraries like
cluster-ensembleexist, but scikit-learn's tools are more than sufficient for most use cases.
内容的提问来源于stack exchange,提问作者IS2057

