降维(PCA/AutoEncoder等)后适用的相似性/相异性度量探讨
Great question—this is something I’ve wrestled with a lot when working with high-dimensional data after dimensionality reduction! Let’s break down the common similarity/dissimilarity metrics you’ll reach for, and when to use each.
欧氏距离(Euclidean Distance): Don’t sleep on this just because you’re in reduced dimensions! If your dimensionality reduction method (like PCA) preserves the original data’s global geometric structure, Euclidean distance still works perfectly. It’s intuitive, easy to compute, and ideal when the reduced space maintains meaningful linear relationships between points.
余弦相似度(Cosine Similarity): This is a go-to when your reduced vectors represent directional relationships rather than absolute magnitude. Think about text data reduced via LSA or embeddings from models like Word2Vec—cosine similarity tells you how aligned two vectors are, ignoring differences in their length. It’s also robust to scaling issues that might creep in during some reduction techniques.
曼哈顿距离(Manhattan Distance): Perfect for reduced spaces where your data is sparse (e.g., after sparse PCA) or you want to treat differences across all dimensions equally (no squaring of large differences like Euclidean does). It’s computationally cheaper than Euclidean in some cases and works well when the reduced space has discrete or grid-like structures.
马氏距离(Mahalanobis Distance): Use this when your reduced dimensions still have correlated features. It accounts for the covariance structure of the reduced data, so it’s more robust than Euclidean distance if the reduced space has non-uniform variance across dimensions. This is especially useful after techniques like Factor Analysis, where dimensions are designed to capture correlated variance.
余弦距离(Cosine Distance): Just
1 - cosine_similarity, this turns the similarity metric into a dissimilarity measure. It’s handy when you need a distance value for algorithms like k-means clustering, instead of a similarity score.
- Use Euclidean if your reduction preserves global linear structure (PCA, MDS).
- Reach for Cosine when working with embeddings/text data where direction matters more than magnitude.
- Go with Manhattan for sparse reduced data or when you want equal weight on all dimension differences.
- Choose Mahalanobis if your reduced dimensions have correlated variance.
内容的提问来源于stack exchange,提问作者KyBe

