如何用Gensim预训练模型聚类词汇与短语?求解短语查询报错
Great question! Let's tackle both your core questions—whether pre-trained models like GoogleNews/Wikipedia work for your clustering task, and how to fix that "not in vocabulary" error—with practical, actionable steps.
1. 能不能用GoogleNews/Wikipedia预训练模型实现聚类?
Absolutely! These models are perfect for this kind of semantic clustering task, even with your limited set of terms. The catch is handling multi-word phrases, since pre-trained word2vec models like GoogleNews don’t include every possible phrase in their vocabulary. But we can work around that easily.
2. 修复"[phrase] is not in the vocabulary"错误
The error pops up because GoogleNews’ vocabulary only includes phrases that were frequent enough in its training corpus to be tokenized as single units (with underscores, like airplane_mode). Phrases like computer_software or some of your loom-related terms might not have made the cut. Here are your most reliable fixes:
Option 1: Average word vectors for missing phrases
For multi-word phrases that aren’t in the vocabulary, calculate the average of the individual word vectors. This gives you a solid approximation of the phrase’s semantic meaning.
Example code tailored to your loom terms:
import gensim import numpy as np # Load the GoogleNews model (as you did before) GOOGLE_MODEL = '../GoogleNews-vectors-negative300.bin' model = gensim.models.KeyedVectors.load_word2vec_format(GOOGLE_MODEL, binary=True) def get_phrase_vector(phrase): words = phrase.split() # Filter out any words not present in the model (unlikely for your terms) valid_vectors = [model[word] for word in words if word in model] if not valid_vectors: raise ValueError(f"No valid words found in phrase: {phrase}") # Return the average vector of the valid words return np.mean(valid_vectors, axis=0) # Test with your phrase print(get_phrase_vector("knit loom")) # Works even if "knit_loom" isn't in the vocab
Option 2: Switch to FastText
FastText is built to handle out-of-vocabulary (OOV) terms by using subword information. Even if a full phrase isn’t in the model, it can generate a vector by combining subword vectors. You can use pre-trained FastText models (like the English Wikipedia variant) with Gensim just like you do with Word2Vec—no extra hassle for your phrase issues.
Option 3: Custom phrase tokens (not needed for your small dataset)
If you had a larger corpus, you could use Gensim’s Phrases module to learn phrase patterns, but since you only have your target terms, this is overkill. Option 1 will work perfectly for your use case.
3. 对针织相关词汇进行聚类
Once you can get vectors for all your terms, you can use a clustering algorithm like K-Means to group semantically similar terms. Here’s a complete, ready-to-run example:
from sklearn.cluster import KMeans # Your target terms list terms = [ "knitting", "knit loom", "loom knitting", "weaving loom", "rainbow loom", "home decoration accessories", "loom knit", "knitting loom" ] # Generate vectors for each term term_vectors = [] valid_terms = [] for term in terms: try: if " " in term: vec = get_phrase_vector(term) else: vec = model[term] term_vectors.append(vec) valid_terms.append(term) except ValueError as e: print(f"Skipping term: {term} (reason: {e})") # Run K-Means clustering (we’ll use 4 clusters based on obvious semantic groups) kmeans = KMeans(n_clusters=4, random_state=42) clusters = kmeans.fit_predict(term_vectors) # Print organized cluster results print("Clustering Results:") for cluster_id in range(kmeans.n_clusters): cluster_terms = [valid_terms[i] for i, c in enumerate(clusters) if c == cluster_id] print(f"Cluster {cluster_id + 1}: {', '.join(cluster_terms)}")
Expected Semantic Clusters (rough output):
- Cluster 1: knitting, weaving loom
- Cluster 2: knit loom, loom knitting, loom knit, knitting loom
- Cluster 3: rainbow loom
- Cluster 4: home decoration accessories
This grouping aligns perfectly with the actual meaning of your terms—core knitting/weaving terms, loom-specific knitting phrases, the niche rainbow loom product, and unrelated home decor accessories.
Final Notes
- GoogleNews works great for general semantic similarity, but if you want hyper-specific craft/knitting vectors, you could train a small model on free knitting text from Wikipedia or craft blogs. That’s extra work, though, and not necessary for your current task.
- FastText is a fantastic upgrade if you want to avoid manually averaging vectors for OOV phrases—it handles them automatically.
内容的提问来源于stack exchange,提问作者CoderTom

