You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于负欧氏距离的谱聚类:为何少见直接使用负距离矩阵谱?

Why Gaussian Kernel Affinity Matrices Are Preferred Over Negative Distance Matrices in Spectral Clustering

Great question—this cuts to the core of why certain design choices become standard in spectral clustering research. Let’s break down the key reasons researchers almost always opt for Gaussian kernel affinity matrices instead of raw negative distance (or squared distance) matrices, even when focusing on the matrix’s spectrum itself:

  • Meaningful similarity scaling
    A negative distance matrix uses values like -d(i,j), which become increasingly negative as points move apart. This gives distant points outsized negative weight that can obscure local cluster structure. In contrast, the Gaussian kernel exp(-d(i,j)²/(2σ²)) maps distances to a bounded (0,1] range, automatically prioritizing nearby points with high similarity weights and rapidly dampening the influence of faraway points. This "soft thresholding" aligns perfectly with clustering’s goal of identifying dense local groups.

  • Positive Semi-Definiteness (PSD) guarantees
    Gaussian kernel matrices are always positive semi-definite (or positive definite, depending on data). This property is critical for spectral analysis: PSD matrices have non-negative eigenvalues, making spectral decomposition stable and interpretable. Most theoretical foundations of spectral clustering (like convergence proofs or consistency guarantees in von Luxburg’s tutorial) rely on PSD matrices. Negative distance/squared distance matrices, however, are rarely PSD—they often have mixed positive and negative eigenvalues, which breaks existing theory and makes spectral results hard to justify.

  • Alignment with manifold learning and kernel method frameworks
    Spectral clustering draws heavily from manifold learning (e.g., Laplacian Eigenmaps) and kernel methods. The Gaussian kernel implicitly maps data into a high-dimensional feature space where similarity metrics align with intuitive notions of "proximity." Negative distance matrices lack this kernel mapping context, so they can’t leverage the rich theoretical toolkit developed for kernel-based methods. This makes it much harder to prove the validity of clustering results, a key priority in academic research.

  • Numerical stability
    Negative distance matrices can have extremely large dynamic ranges (especially in high-dimensional data), leading to poor condition numbers and numerical instability during spectral decomposition. Gaussian kernels compress all similarity values into a narrow (0,1] range, which improves numerical stability and ensures reliable results in practice.

  • Historical standardization
    Early foundational work in spectral clustering (including Michael Jordan’s NIPS papers) adopted Gaussian kernels, establishing a common framework for subsequent research. Using this standard makes experimental comparisons easier and allows researchers to build on existing theory, whereas switching to negative distance matrices would require re-deriving core proofs and benchmarks—a high barrier that discourages exploration.

内容的提问来源于stack exchange,提问作者Timothy Chu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 09:13:55