You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Word2Vec计算颜色词Cosine Similarity结果异常,求问题排查及原理验证

问题

我尝试计算句子中两个单词的余弦相似度,句子为“The black cat sat on the couch and the brown dog slept on the rug”。以下是我的Python代码:

from nltk.tokenize import sent_tokenize, word_tokenize
import warnings
 
warnings.filterwarnings(action = 'ignore')
 
import gensim
from gensim.models import Word2Vec
from sklearn.metrics.pairwise import cosine_similarity

sentence = "The black cat sat on the couch and the brown dog slept on the rug"
# Replaces escape character with space
f = sentence.replace("\n", " ")
 
data = []

# sentence parsing
for i in sent_tokenize(f):
    temp = []
    # tokenize the sentence into words
    for j in word_tokenize(i):
        temp.append(j.lower())
    data.append(temp)
print(data)
# Creating Skip Gram model
model2 = gensim.models.Word2Vec(data, min_count = 1, vector_size = 512, window = 5, sg = 1)

# Print results
print("Cosine similarity between 'black' " +
          "and 'brown' - Skip Gram : ",
    model2.wv.similarity('black', 'brown'))

实际输出结果为:

[['the', 'black', 'cat', 'sat', 'on', 'the', 'couch', 'and', 'the', 'brown', 'dog', 'slept', 'on', 'the', 'rug']]
Cosine similarity between 'black' and 'brown' - Skip Gram :  0.008911405

我认为“black”和“brown”均为颜色类词汇,它们的余弦相似度应接近1,但结果远低于预期。请问问题出在哪里?我对余弦相似度的理解是否正确?

问题分析与解决

1. 余弦相似度的理解无误

余弦相似度取值范围为[-1, 1],值越接近1代表两个向量方向越一致,语义相似度越高。“black”与“brown”同属颜色类词汇,理想情况下二者的向量相似度确实应该处于较高水平。

2. 核心问题:训练语料规模过小

Word2Vec这类词嵌入模型的本质是通过大规模上下文语料学习单词的语义关联。仅用单句语料训练时,模型无法捕捉到“black”和“brown”作为颜色词的共性:

  • 在目标句子中,“black”的上下文是the, cat, sat,“brown”的上下文是the, dog, slept,二者上下文完全无重叠,模型无法建立二者的语义关联。
  • 极小语料下,向量初始化的随机噪声会主导最终结果,生成的向量不具备任何语义代表性。

3. 可行解决方案

方案一:使用预训练词嵌入模型

直接加载经过大规模语料训练的成熟模型(如Google Word2Vec预训练模型、GloVe、FastText等),这类模型已学习到通用的单词语义关联。示例代码:

from gensim.models import KeyedVectors

# 加载预训练模型(需替换为本地模型文件路径)
pretrained_model = KeyedVectors.load_word2vec_format('GoogleNews-vectors-negative300.bin', binary=True)
print("Cosine similarity between 'black' and 'brown':", pretrained_model.similarity('black', 'brown'))

运行后会得到符合预期的高相似度值(通常在0.6以上)。

方案二:扩充训练语料

若坚持自行训练,需收集包含大量颜色词上下文的语料(如包含各类颜色描述的文本集合),语料规模至少达到数千句级别,模型才能学习到有意义的语义关系。

内容的提问来源于stack exchange,提问作者sats

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.27 03:26:13