You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

sklearn KMeans中precompute_distances参数作用及预计算距离问询

Understanding precompute_distances in scikit-learn's KMeans

Great question—let's break this down clearly, since this parameter is all about balancing speed and memory usage in your clustering job.

What does precompute_distances do?

At its core, this parameter controls whether KMeans precomputes certain distance-related values to avoid redundant calculations during its iterative process. KMeans spends most of its time calculating distances between samples and cluster centroids to assign points to the nearest cluster. Precomputing these values speeds things up, but it uses more memory—so the parameter lets you tweak this tradeoff.

It accepts three values:

  • 'auto' (default): Scikit-learn automatically decides. If the product of n_samples * n_clusters exceeds 12 million, it skips precomputing (this would eat up ~100MB of memory per job with double-precision calculations). For smaller datasets, it enables precomputing to speed things up.
  • True: Forces precomputing regardless of dataset size. Use this when you have plenty of RAM and want the fastest possible runtime for small-to-medium datasets.
  • False: Disables precomputing entirely. The algorithm recalculates distances from scratch in every iteration, which saves memory but is slower—ideal for large datasets where RAM is limited.

Which distances are precomputed?

The key things being precomputed are:

  1. Squared L2 norms of all sample points: That is, the squared distance of each sample from the origin (||x||² for every sample x). KMeans uses Euclidean distance by default, and we can rewrite squared Euclidean distance between a sample x and centroid c as:
    ||x - c||² = ||x||² + ||c||² - 2 * x·c
    
    Precomputing ||x||² lets us avoid recalculating this part every time we need to compute distances to a new centroid.
  2. Full sample-to-centroid distance matrix (for small datasets): When n_samples * n_clusters is small enough (under 12 million in auto mode), the algorithm precomputes a full matrix where each row is a sample and each column is the distance from that sample to a centroid. This matrix is reused during the iteration steps to avoid recalculating distances multiple times.

In short, it's not precomputing all pairwise sample distances (KMeans doesn't need those!), just the values that make sample-to-centroid distance calculations faster.

内容的提问来源于stack exchange,提问作者user9562553

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:29:23