sklearn KMeans中precompute_distances参数作用及预计算距离问询
precompute_distances in scikit-learn's KMeans Great question—let's break this down clearly, since this parameter is all about balancing speed and memory usage in your clustering job.
What does precompute_distances do?
At its core, this parameter controls whether KMeans precomputes certain distance-related values to avoid redundant calculations during its iterative process. KMeans spends most of its time calculating distances between samples and cluster centroids to assign points to the nearest cluster. Precomputing these values speeds things up, but it uses more memory—so the parameter lets you tweak this tradeoff.
It accepts three values:
'auto'(default): Scikit-learn automatically decides. If the product ofn_samples * n_clustersexceeds 12 million, it skips precomputing (this would eat up ~100MB of memory per job with double-precision calculations). For smaller datasets, it enables precomputing to speed things up.True: Forces precomputing regardless of dataset size. Use this when you have plenty of RAM and want the fastest possible runtime for small-to-medium datasets.False: Disables precomputing entirely. The algorithm recalculates distances from scratch in every iteration, which saves memory but is slower—ideal for large datasets where RAM is limited.
Which distances are precomputed?
The key things being precomputed are:
- Squared L2 norms of all sample points: That is, the squared distance of each sample from the origin (
||x||²for every samplex). KMeans uses Euclidean distance by default, and we can rewrite squared Euclidean distance between a samplexand centroidcas:
Precomputing||x - c||² = ||x||² + ||c||² - 2 * x·c||x||²lets us avoid recalculating this part every time we need to compute distances to a new centroid. - Full sample-to-centroid distance matrix (for small datasets): When
n_samples * n_clustersis small enough (under 12 million in auto mode), the algorithm precomputes a full matrix where each row is a sample and each column is the distance from that sample to a centroid. This matrix is reused during the iteration steps to avoid recalculating distances multiple times.
In short, it's not precomputing all pairwise sample distances (KMeans doesn't need those!), just the values that make sample-to-centroid distance calculations faster.
内容的提问来源于stack exchange,提问作者user9562553

