Python无参聚类方案咨询:仅输入数据自动确定簇数与欧氏距离阈值
Hey there! Great question—yes, you absolutely can run parameter-free clustering in Python where the algorithm automatically figures out both the number of clusters and the Euclidean distance threshold straight from your dataset. Let’s dive into the best solutions that fit your exact needs:
Top Recommendations
1. HDBSCAN (Hierarchical DBSCAN)
This is hands down the most fitting option for what you’re asking for. HDBSCAN eliminates the need to manually set a distance threshold (eps) or cluster count entirely. Instead, it works by analyzing the density hierarchy of your data to identify stable, meaningful clusters on its own, plus it automatically flags noise points that don’t belong to any cluster.
How it works
HDBSCAN builds a tree of clusters based on data density, then picks out the clusters that stay consistent across different density levels (these are the "stable" ones). It essentially learns the optimal distance threshold by looking at how persistent clusters are as you adjust density cutoffs.
Code Example
import hdbscan from sklearn.datasets import make_blobs # Replace this with your actual dataset X, _ = make_blobs(n_samples=500, centers=4, random_state=42) # Initialize HDBSCAN - no eps or n_clusters needed! # min_cluster_size is optional (filters tiny, irrelevant clusters) clusterer = hdbscan.HDBSCAN(min_cluster_size=5) cluster_labels = clusterer.fit_predict(X) # Check results unique_clusters = set(cluster_labels) # Subtract 1 if noise points (-1) are present cluster_count = len(unique_clusters) - (1 if -1 in unique_clusters else 0) print(f"Automatically identified cluster count: {cluster_count}") print(f"Cluster labels ( -1 = noise ): {cluster_labels}")
2. DBSCAN with Automatic eps Selection
DBSCAN is a tried-and-true density-based algorithm, but normally you’d need to set eps (the distance threshold) and min_samples (minimum points for a dense region). The good news is you can automate the eps selection using the k-distance graph trick—no guesswork required.
How it works
We calculate the distance from each data point to its k-th nearest neighbor (where k matches min_samples). When we sort these distances and plot them, the "elbow" point (where distances jump sharply) tells us the optimal eps: this is the line between dense cluster regions and sparse noise. You can even automate finding that elbow instead of eyeballing the plot.
Code Example
import numpy as np import matplotlib.pyplot as plt from sklearn.cluster import DBSCAN from sklearn.neighbors import NearestNeighbors from sklearn.datasets import make_blobs # Replace with your dataset X, _ = make_blobs(n_samples=500, centers=4, random_state=42) # Step 1: Calculate k-distances to find optimal eps min_samples = 5 neighbors = NearestNeighbors(n_neighbors=min_samples) neighbors_fit = neighbors.fit(X) distances, _ = neighbors_fit.kneighbors(X) # Sort distances and find the elbow point distances = np.sort(distances[:, min_samples-1], axis=0) def auto_find_elbow(distances): # Calculate second derivative to detect the sharp jump (elbow) dx = np.diff(np.arange(len(distances))) dy = np.diff(distances) second_deriv = np.diff(dy)/dx[:-1] # The elbow is where the second derivative is maximum elbow_idx = np.argmax(second_deriv) + 1 return distances[elbow_idx] eps = auto_find_elbow(distances) print(f"Automatically determined eps threshold: {eps:.4f}") # Optional: Plot the k-distance graph to verify plt.plot(distances) plt.title('K-Distance Graph (Elbow = Optimal eps)') plt.xlabel('Data Points Sorted by Distance') plt.ylabel('Distance to k-th Nearest Neighbor') plt.show() # Step 2: Run DBSCAN with auto-detected eps dbscan = DBSCAN(eps=eps, min_samples=min_samples) cluster_labels = dbscan.fit_predict(X) unique_clusters = set(cluster_labels) cluster_count = len(unique_clusters) - (1 if -1 in unique_clusters else 0) print(f"Automatically identified cluster count: {cluster_count}")
3. OPTICS (Ordering Points To Identify the Clustering Structure)
OPTICS is another density-based algorithm that skips needing a fixed eps value. It creates a reachability plot that maps out the density structure of your data, and you can set it to automatically extract clusters from this plot without manual input.
Code Example
from sklearn.cluster import OPTICS from sklearn.datasets import make_blobs # Replace with your dataset X, _ = make_blobs(n_samples=500, centers=4, random_state=42) # Initialize OPTICS - no fixed eps needed # cluster_method='xi' tells it to auto-extract clusters from reachability plot optics = OPTICS(min_samples=5, cluster_method='xi') cluster_labels = optics.fit_predict(X) unique_clusters = set(cluster_labels) cluster_count = len(unique_clusters) - (1 if -1 in unique_clusters else 0) print(f"Automatically identified cluster count: {cluster_count}")
Which One Should You Pick?
- HDBSCAN is my top recommendation. It’s robust to clusters of varying sizes and densities, requires almost no manual tuning, and perfectly handles both auto-detecting cluster counts and distance thresholds. It’s ideal for most real-world datasets.
- DBSCAN with auto
epsis a solid choice if you prefer a classic, well-understood algorithm, though it needs a bit extra code to automate theepsselection. - OPTICS is great if you want to explore the full density hierarchy of your data, but it can be slower on large datasets compared to HDBSCAN.
内容的提问来源于stack exchange,提问作者Utpal Datta

