You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

300维嵌入转2维:如何降低聚类重叠优化句子分组可视化?

Optimizing 2D Visualization of Sentence Embeddings with Reduced Cluster Overlap

Hey there! Let’s work through this problem—getting clean, well-separated clusters when reducing your 300-dimensional sentence embeddings (after mean pooling) down to 2D. You’ve already tried the basics (raw T-SNE and PCA→T-SNE), so let’s dive into actionable tweaks and alternative approaches that’ll help with cluster separation.

1. Fine-Tune the PCA + T-SNE Pipeline (Most Straightforward Win)

The PCA→T-SNE combo is a solid start, but default parameters often aren’t optimized for text embeddings. Here’s how to tweak it:

Step 1: Standardize Your Embeddings First

Text embeddings can have varying scales across dimensions, which throws off PCA and T-SNE. Start by normalizing your data:

from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.manifold import TSNE
import numpy as np

# Assume your mean-pooled embeddings are stored in a numpy array `embeddings` (shape: [num_samples, 300])
scaler = StandardScaler()
scaled_embeddings = scaler.fit_transform(embeddings)

Step 2: Optimize T-SNE Hyperparameters

T-SNE’s perplexity and learning_rate are make-or-break for cluster separation:

  • Perplexity: Should be roughly sqrt(num_samples) (e.g., 30-50 for 1000 samples). It balances local and global structure.
  • Learning Rate: If clusters look squished, try increasing it (300-1000); if they’re too spread out, drop it to 100-200.
  • n_iter: Increase to 5000+ to ensure the algorithm converges fully.

Example code:

# First reduce to 50D with PCA
pca = PCA(n_components=50)
pca_embeddings = pca.fit_transform(scaled_embeddings)

# Then run T-SNE with optimized params
tsne = TSNE(
    n_components=2,
    perplexity=40,  # Adjust based on your sample size
    learning_rate=300,
    n_iter=5000,
    random_state=42
)
tsne_2d = tsne.fit_transform(pca_embeddings)

2. Switch to UMAP for Better Global Structure Preservation

UMAP often outperforms T-SNE at retaining global cluster relationships, which can lead to more separated groups. It’s also faster for large datasets.

Key UMAP parameters to tweak:

  • n_neighbors: Controls how much local vs global structure is preserved (smaller = more local; try 10-30 for text).
  • min_dist: Determines how tightly clusters are packed (lower values = tighter, more separated clusters).

Example code:

import umap

# You can skip PCA if you want, but UMAP handles 300D well too
umap_2d = umap.UMAP(
    n_components=2,
    n_neighbors=20,
    min_dist=0.1,
    random_state=42
).fit_transform(scaled_embeddings)

3. Add Supervision (If You Have Cluster Labels)

If you already have ground-truth labels for your sentence groups, use supervised or semi-supervised methods to force separation:

Option A: LDA + T-SNE

Linear Discriminant Analysis (LDA) maximizes separation between known classes. Combine it with T-SNE for better 2D visualization:

from sklearn.discriminant_analysis import LinearDiscriminantAnalysis

# Assume `labels` is your array of ground-truth cluster labels
lda = LinearDiscriminantAnalysis(n_components=2)  # LDA can only go up to num_classes-1
lda_embeddings = lda.fit_transform(scaled_embeddings, labels)

# If you have more than 3 classes, run T-SNE on LDA output:
# lda_tsne = TSNE(n_components=2, random_state=42).fit_transform(lda_embeddings)

Option B: Supervised UMAP

UMAP supports supervised learning by passing labels to the y parameter:

supervised_umap = umap.UMAP(
    n_components=2,
    n_neighbors=20,
    min_dist=0.1,
    random_state=42
).fit_transform(scaled_embeddings, y=labels)

4. Preprocess Embeddings for Better Signal

  • L2 Normalization: For sentence embeddings (like BERT, Sentence-BERT), L2-normalizing each vector ensures all embeddings have the same magnitude, which reduces noise in downstream steps:
    normalized_embeddings = embeddings / np.linalg.norm(embeddings, axis=1, keepdims=True)
    
  • Feature Selection: Remove low-variance dimensions (they add noise without signal):
    from sklearn.feature_selection import VarianceThreshold
    selector = VarianceThreshold(threshold=0.01)  # Adjust threshold based on your data
    filtered_embeddings = selector.fit_transform(scaled_embeddings)
    

Final Tips

  • Always visualize multiple parameter combinations—what works for one dataset might not work for another.
  • If clusters still overlap, try increasing the PCA dimension (e.g., 100 instead of 50) to retain more signal before T-SNE/UMAP.

内容的提问来源于stack exchange,提问作者Tadeusz Jasińska

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:50:38