You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python无参聚类方案咨询:仅输入数据自动确定簇数与欧氏距离阈值

Hey there! Great question—yes, you absolutely can run parameter-free clustering in Python where the algorithm automatically figures out both the number of clusters and the Euclidean distance threshold straight from your dataset. Let’s dive into the best solutions that fit your exact needs:

Parameter-Free Clustering in Python: Auto-Determine Clusters & Distance Thresholds

Top Recommendations

1. HDBSCAN (Hierarchical DBSCAN)

This is hands down the most fitting option for what you’re asking for. HDBSCAN eliminates the need to manually set a distance threshold (eps) or cluster count entirely. Instead, it works by analyzing the density hierarchy of your data to identify stable, meaningful clusters on its own, plus it automatically flags noise points that don’t belong to any cluster.

How it works

HDBSCAN builds a tree of clusters based on data density, then picks out the clusters that stay consistent across different density levels (these are the "stable" ones). It essentially learns the optimal distance threshold by looking at how persistent clusters are as you adjust density cutoffs.

Code Example

import hdbscan
from sklearn.datasets import make_blobs

# Replace this with your actual dataset
X, _ = make_blobs(n_samples=500, centers=4, random_state=42)

# Initialize HDBSCAN - no eps or n_clusters needed!
# min_cluster_size is optional (filters tiny, irrelevant clusters)
clusterer = hdbscan.HDBSCAN(min_cluster_size=5)
cluster_labels = clusterer.fit_predict(X)

# Check results
unique_clusters = set(cluster_labels)
# Subtract 1 if noise points (-1) are present
cluster_count = len(unique_clusters) - (1 if -1 in unique_clusters else 0)
print(f"Automatically identified cluster count: {cluster_count}")
print(f"Cluster labels ( -1 = noise ): {cluster_labels}")

2. DBSCAN with Automatic eps Selection

DBSCAN is a tried-and-true density-based algorithm, but normally you’d need to set eps (the distance threshold) and min_samples (minimum points for a dense region). The good news is you can automate the eps selection using the k-distance graph trick—no guesswork required.

How it works

We calculate the distance from each data point to its k-th nearest neighbor (where k matches min_samples). When we sort these distances and plot them, the "elbow" point (where distances jump sharply) tells us the optimal eps: this is the line between dense cluster regions and sparse noise. You can even automate finding that elbow instead of eyeballing the plot.

Code Example

import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import DBSCAN
from sklearn.neighbors import NearestNeighbors
from sklearn.datasets import make_blobs

# Replace with your dataset
X, _ = make_blobs(n_samples=500, centers=4, random_state=42)

# Step 1: Calculate k-distances to find optimal eps
min_samples = 5
neighbors = NearestNeighbors(n_neighbors=min_samples)
neighbors_fit = neighbors.fit(X)
distances, _ = neighbors_fit.kneighbors(X)

# Sort distances and find the elbow point
distances = np.sort(distances[:, min_samples-1], axis=0)

def auto_find_elbow(distances):
    # Calculate second derivative to detect the sharp jump (elbow)
    dx = np.diff(np.arange(len(distances)))
    dy = np.diff(distances)
    second_deriv = np.diff(dy)/dx[:-1]
    # The elbow is where the second derivative is maximum
    elbow_idx = np.argmax(second_deriv) + 1
    return distances[elbow_idx]

eps = auto_find_elbow(distances)
print(f"Automatically determined eps threshold: {eps:.4f}")

# Optional: Plot the k-distance graph to verify
plt.plot(distances)
plt.title('K-Distance Graph (Elbow = Optimal eps)')
plt.xlabel('Data Points Sorted by Distance')
plt.ylabel('Distance to k-th Nearest Neighbor')
plt.show()

# Step 2: Run DBSCAN with auto-detected eps
dbscan = DBSCAN(eps=eps, min_samples=min_samples)
cluster_labels = dbscan.fit_predict(X)

unique_clusters = set(cluster_labels)
cluster_count = len(unique_clusters) - (1 if -1 in unique_clusters else 0)
print(f"Automatically identified cluster count: {cluster_count}")

3. OPTICS (Ordering Points To Identify the Clustering Structure)

OPTICS is another density-based algorithm that skips needing a fixed eps value. It creates a reachability plot that maps out the density structure of your data, and you can set it to automatically extract clusters from this plot without manual input.

Code Example

from sklearn.cluster import OPTICS
from sklearn.datasets import make_blobs

# Replace with your dataset
X, _ = make_blobs(n_samples=500, centers=4, random_state=42)

# Initialize OPTICS - no fixed eps needed
# cluster_method='xi' tells it to auto-extract clusters from reachability plot
optics = OPTICS(min_samples=5, cluster_method='xi')
cluster_labels = optics.fit_predict(X)

unique_clusters = set(cluster_labels)
cluster_count = len(unique_clusters) - (1 if -1 in unique_clusters else 0)
print(f"Automatically identified cluster count: {cluster_count}")

Which One Should You Pick?

  • HDBSCAN is my top recommendation. It’s robust to clusters of varying sizes and densities, requires almost no manual tuning, and perfectly handles both auto-detecting cluster counts and distance thresholds. It’s ideal for most real-world datasets.
  • DBSCAN with auto eps is a solid choice if you prefer a classic, well-understood algorithm, though it needs a bit extra code to automate the eps selection.
  • OPTICS is great if you want to explore the full density hierarchy of your data, but it can be slower on large datasets compared to HDBSCAN.

内容的提问来源于stack exchange,提问作者Utpal Datta

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:18:35