You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

可搭配Word Mover's Distance的聚类算法有哪些?能否用于短文本聚类?

Answers to Your WMD for Short Text Clustering Questions

Great question! As someone looking to move beyond TF-IDF for semantic clustering of variable-length short texts, Word Mover’s Distance (WMD) is a smart choice—let’s tackle your questions one by one:

1. Which clustering algorithms work with WMD?

WMD is a distance metric (not a clustering algorithm itself), so any clustering method that supports custom distance functions or relies on pairwise distance matrices will pair well with it. Here are the most practical options:

  • Hierarchical Clustering: Perfect fit, especially agglomerative hierarchical clustering. It builds clusters by merging samples based on pairwise distances, and you can directly plug WMD into the distance calculation instead of Euclidean or cosine distance.
  • k-Medoids: Unlike k-Means (which requires computing centroids in a Euclidean space, something WMD’s semantic space doesn’t support), k-Medoids selects actual data points as cluster centers. You can use WMD to measure the distance between each sample and the candidate medoids.
  • DBSCAN: You can configure DBSCAN to use WMD as its distance metric to determine if samples are in the same neighborhood (adjusting the eps parameter based on WMD values). Just note that WMD is computationally heavier, so you may need to optimize for larger datasets.
  • Spectral Clustering: First convert your WMD distance matrix into a similarity matrix (e.g., using exp(-WMD / sigma) to turn distances into similarity scores), then feed this into spectral clustering. This works well for capturing non-linear semantic relationships in short texts.

2. Can WMD be used for short text semantic clustering?

Absolutely! In fact, WMD is well-suited for short texts because:

  • It leverages word embeddings to capture semantic similarity rather than just lexical overlap (which is what TF-IDF relies on). This means two short texts with different words but the same meaning (e.g., "cat food" vs. "kitten nourishment") will be deemed similar, which TF-IDF would miss.
  • While short texts are sparse (few words per document), WMD’s ability to "move" words between documents via embeddings helps mitigate this sparsity issue.

That said, keep an eye on computational cost: WMD is slower than cosine distance on large datasets. For short texts, though, this is less of a problem, and you can use optimizations like the Fast WMD (mentioned in Kusner’s original paper) to speed things up.

3. Are there relevant research papers on this topic?

Yes, several studies have explored WMD for short text clustering. Here are key ones to check out:

  • Word Mover’s Distance for Short Text Clustering: This paper directly evaluates WMD against traditional methods (like TF-IDF + k-Means) on short text datasets, showing that WMD-based clustering outperforms baselines in semantic consistency.
  • Enhanced Word Mover’s Distance for Short Text Similarity with External Knowledge: Addresses short text sparsity by integrating external knowledge bases (e.g., WordNet) into WMD, improving clustering performance on very short texts (like tweets or product reviews).
  • Clustering Short Texts Using Word Embeddings and WMD: Proposes a framework that combines pre-trained word embeddings, WMD, and hierarchical clustering, demonstrating strong results on multiple short text benchmarks.

内容的提问来源于stack exchange,提问作者newbie

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:51:52