Gensim most_similar工作原理及两种调用结果差异原因问询
most_similar Behaves Differently Between Your Two Call Styles Let’s break down exactly what’s going on here—because these two approaches aren’t doing the same thing under the hood, even though they seem like they should.
1. The positive/negative Parameter Path Uses Structured Normalization
When you call model.most_similar(positive=['woman', 'king'], negative=['man']), Gensim follows a precise, opinionated workflow:
- For every word in
positiveandnegative, it first pulls the L2-normalized version of their vectors (scaled to a length of 1, so only direction matters). - It sums all normalized positive vectors, subtracts all normalized negative vectors to create a raw target vector.
- It then re-normalizes this combined target vector to ensure it’s also a unit vector.
- Finally, it calculates cosine similarity between this normalized target and all normalized word vectors in the model, and returns the top results—automatically excluding the input words (
woman,king,man) from the output list.
2. Manual Vector Math Skips Per-Word Normalization
When you compute model["king"] + model["woman"] - model["man"] and pass that to most_similar, you’re bypassing the critical per-word normalization step:
- You’re adding/subtracting the raw, unnormalized vectors directly. These vectors have varying lengths (not unit length) depending on how the model trained.
most_similarwill still normalize your final combined vector before calculating similarities, but starting from raw vectors vs per-word normalized vectors creates a fundamentally different target vector.- Additionally, since you’re passing a custom vector instead of using the
positive/negativeparameters, Gensim has no way to know which words to exclude—soking(which has a high similarity to your manually computed vector) shows up in the results.
3. Why Similarity Scores & Rankings Differ
Cosine similarity depends entirely on the direction of vectors. Since:
- The built-in path builds a target vector from normalized components, then re-normalizes.
- Your manual path builds a target vector from raw components, then normalizes.
These two target vectors point in slightly different directions. When compared against the model’s normalized word vectors, their cosine similarities (and thus result rankings) will differ—including the score for queen.
Replicating the Built-In Behavior Manually
If you want to match the output of the positive/negative call with manual math, mirror Gensim’s steps:
import numpy as np # Get normalized vectors for each input word norm_king = model.wv.get_vector('king', norm=True) norm_woman = model.wv.get_vector('woman', norm=True) norm_man = model.wv.get_vector('man', norm=True) # Compute target vector (sum positives, subtract negatives) target = norm_king + norm_woman - norm_man # Re-normalize the target vector target = target / np.linalg.norm(target) # Get top results, then manually exclude input words results = model.wv.most_similar([target], topn=10) filtered_results = [(word, score) for word, score in results if word not in {'king', 'woman', 'man'}]
This will give you results nearly identical to the built-in most_similar call.
内容的提问来源于stack exchange,提问作者Peyman

