You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Gensim Word2Vec中most_similar与similar_by_vector的结果差异原因

Why most_similar and similar_by_vector Return Different Results for Vector Operations?

Great question—this is a common gotcha with Gensim's Word2Vec implementations, and the difference boils down to two key behaviors in how each method processes inputs:

1. most_similar automatically excludes input words from results

By default, most_similar filters out any words you pass in the positive or negative parameters from its returned top matches. In your example, that means king, man, and woman are all excluded from the output.

Compare that to similar_by_vector: it doesn't do any automatic filtering. That's why you see ('king', 0.8655...) as the first result in your similar_by_vector call—since you're passing a vector derived from king, the model correctly identifies king as the closest match, and there's no rule stopping it from showing up.

2. most_similar normalizes vectors during combination

Under the hood, most_similar doesn't just add/subtract raw vectors like you did manually. Here's what it actually does when you pass positive=['king', 'woman'], negative=['man']:

  • It first L2-normalizes each individual word vector (scales them to have a length of 1).
  • It sums the normalized positive vectors, subtracts the normalized negative vectors to create the target query vector.
  • It then L2-normalizes this final query vector before computing similarities.

When you calculate q = model['king'] - model['man'] + model['woman'] manually, you're using the raw, unnormalized vectors (unless you explicitly normalized them first). This creates a query vector that's mathematically different from the one most_similar generates, leading to different cosine similarity scores for matching words (like the 0.75 vs 0.767 for queen).

If you want to replicate most_similar's result with similar_by_vector, you'd need to:

  • Normalize each input vector before combining them
  • Normalize the final combined vector
  • Exclude the input words (king, man, woman) from the results

For example:

from gensim.matutils import unitvec

# Normalize each vector first
king_norm = unitvec(model['king'])
man_norm = unitvec(model['man'])
woman_norm = unitvec(model['woman'])

# Create normalized query vector
q = unitvec(king_norm - man_norm + woman_norm)

# Get results, excluding input words
results = model.similar_by_vector(q, topn=10, restrict_vocab=None)
filtered_results = [(word, score) for word, score in results if word not in {'king', 'man', 'woman'}]
print(filtered_results)

This should give you results nearly identical to model.most_similar(positive=['king', 'woman'], negative=['man']).

内容的提问来源于stack exchange,提问作者peidaqi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 07:01:34