可搭配Word Mover's Distance的聚类算法有哪些?能否用于短文本聚类?
Great question! As someone looking to move beyond TF-IDF for semantic clustering of variable-length short texts, Word Mover’s Distance (WMD) is a smart choice—let’s tackle your questions one by one:
1. Which clustering algorithms work with WMD?
WMD is a distance metric (not a clustering algorithm itself), so any clustering method that supports custom distance functions or relies on pairwise distance matrices will pair well with it. Here are the most practical options:
- Hierarchical Clustering: Perfect fit, especially agglomerative hierarchical clustering. It builds clusters by merging samples based on pairwise distances, and you can directly plug WMD into the distance calculation instead of Euclidean or cosine distance.
- k-Medoids: Unlike k-Means (which requires computing centroids in a Euclidean space, something WMD’s semantic space doesn’t support), k-Medoids selects actual data points as cluster centers. You can use WMD to measure the distance between each sample and the candidate medoids.
- DBSCAN: You can configure DBSCAN to use WMD as its distance metric to determine if samples are in the same neighborhood (adjusting the
epsparameter based on WMD values). Just note that WMD is computationally heavier, so you may need to optimize for larger datasets. - Spectral Clustering: First convert your WMD distance matrix into a similarity matrix (e.g., using
exp(-WMD / sigma)to turn distances into similarity scores), then feed this into spectral clustering. This works well for capturing non-linear semantic relationships in short texts.
2. Can WMD be used for short text semantic clustering?
Absolutely! In fact, WMD is well-suited for short texts because:
- It leverages word embeddings to capture semantic similarity rather than just lexical overlap (which is what TF-IDF relies on). This means two short texts with different words but the same meaning (e.g., "cat food" vs. "kitten nourishment") will be deemed similar, which TF-IDF would miss.
- While short texts are sparse (few words per document), WMD’s ability to "move" words between documents via embeddings helps mitigate this sparsity issue.
That said, keep an eye on computational cost: WMD is slower than cosine distance on large datasets. For short texts, though, this is less of a problem, and you can use optimizations like the Fast WMD (mentioned in Kusner’s original paper) to speed things up.
3. Are there relevant research papers on this topic?
Yes, several studies have explored WMD for short text clustering. Here are key ones to check out:
- Word Mover’s Distance for Short Text Clustering: This paper directly evaluates WMD against traditional methods (like TF-IDF + k-Means) on short text datasets, showing that WMD-based clustering outperforms baselines in semantic consistency.
- Enhanced Word Mover’s Distance for Short Text Similarity with External Knowledge: Addresses short text sparsity by integrating external knowledge bases (e.g., WordNet) into WMD, improving clustering performance on very short texts (like tweets or product reviews).
- Clustering Short Texts Using Word Embeddings and WMD: Proposes a framework that combines pre-trained word embeddings, WMD, and hierarchical clustering, demonstrating strong results on multiple short text benchmarks.
内容的提问来源于stack exchange,提问作者newbie

