Gensim Word2Vec中most_similar与similar_by_vector的结果差异原因
most_similar and similar_by_vector Return Different Results for Vector Operations? Great question—this is a common gotcha with Gensim's Word2Vec implementations, and the difference boils down to two key behaviors in how each method processes inputs:
1. most_similar automatically excludes input words from results
By default, most_similar filters out any words you pass in the positive or negative parameters from its returned top matches. In your example, that means king, man, and woman are all excluded from the output.
Compare that to similar_by_vector: it doesn't do any automatic filtering. That's why you see ('king', 0.8655...) as the first result in your similar_by_vector call—since you're passing a vector derived from king, the model correctly identifies king as the closest match, and there's no rule stopping it from showing up.
2. most_similar normalizes vectors during combination
Under the hood, most_similar doesn't just add/subtract raw vectors like you did manually. Here's what it actually does when you pass positive=['king', 'woman'], negative=['man']:
- It first L2-normalizes each individual word vector (scales them to have a length of 1).
- It sums the normalized positive vectors, subtracts the normalized negative vectors to create the target query vector.
- It then L2-normalizes this final query vector before computing similarities.
When you calculate q = model['king'] - model['man'] + model['woman'] manually, you're using the raw, unnormalized vectors (unless you explicitly normalized them first). This creates a query vector that's mathematically different from the one most_similar generates, leading to different cosine similarity scores for matching words (like the 0.75 vs 0.767 for queen).
If you want to replicate most_similar's result with similar_by_vector, you'd need to:
- Normalize each input vector before combining them
- Normalize the final combined vector
- Exclude the input words (
king,man,woman) from the results
For example:
from gensim.matutils import unitvec # Normalize each vector first king_norm = unitvec(model['king']) man_norm = unitvec(model['man']) woman_norm = unitvec(model['woman']) # Create normalized query vector q = unitvec(king_norm - man_norm + woman_norm) # Get results, excluding input words results = model.similar_by_vector(q, topn=10, restrict_vocab=None) filtered_results = [(word, score) for word, score in results if word not in {'king', 'man', 'woman'}] print(filtered_results)
This should give you results nearly identical to model.most_similar(positive=['king', 'woman'], negative=['man']).
内容的提问来源于stack exchange,提问作者peidaqi

