You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于词向量欧氏距离矩阵获取指定词的n个最相似词

问题背景

我使用Keras手动计算得到词向量矩阵word_embeddings,矩阵共1000行25列,行索引为对应词汇,列对应词向量各维度取值,结构示例如下:

>>> word_embeddings

        0           1           2           3 
movie   0.007964    0.004251    -0.049078   0.032954    ...
film    -0.006703   0.045888    -0.020975   0.012483    ...
one     -0.011733   0.003348    -0.022017   -0.006476   ...
make    0.045888    -0.011219   0.037796    -0.041868   ...

1000 rows × 25 columns

需求为:给定输入词,获取与其最相似的n个词,例如输入input='movie'时,输出结果为['film', 'cinema', ...]。
目前已通过sklearn的euclidean_distances方法计算得到欧氏距离矩阵,需要基于该矩阵实现上述需求,距离矩阵计算代码及结果如下:

>>> from sklearn.metrics.pairwise import euclidean_distances
>>> distance_matrix = euclidean_distances(word_embeddings)

array([[0.       , 2.4705646, 2.363872 , ..., 3.1345532, 2.9737253,
        2.791427 ],
       [2.4705646, 0.       , 2.3540049, ..., 3.6580865, 3.4589343,
        3.494087 ],
       [2.363872 , 2.3540049, 0.       , ..., 3.9583569, 3.692863 ,
        3.5237448],
       ...,
       [3.1345532, 3.6580865, 3.9583569, ..., 0.       , 4.0572405,
        4.0648513],
       [2.9737253, 3.4589343, 3.692863 , ..., 4.0572405, 0.       ,
        4.156624 ],
       [2.791427 , 3.494087 , 3.5237448, ..., 4.0648513, 4.156624 ,
        0.       ]], dtype=float32)

1000 rows × 1000 columns
实现方案

欧氏距离越小,两个词的语义相似度越高,按以下步骤处理即可:

  • 先存储词表的索引映射,实现词汇和矩阵行号的互相转换
  • 定位输入词在距离矩阵中对应的行,取出该行所有距离值
  • 对距离值做升序排序,跳过第一个值(词与自身的距离,恒为0),取前n个最小距离对应的行号
  • 将行号映射回对应词汇,即为最终的相似词结果

直接复用以下代码即可:

import numpy as np

def get_topn_similar_words(input_word, word_embeddings, distance_matrix, topn=10):
    # 构建词和位置的映射
    word_list = word_embeddings.index.to_list()
    word2idx = {word: idx for idx, word in enumerate(word_list)}
    
    if input_word not in word2idx:
        raise ValueError(f"输入词 {input_word} 不在当前词表范围内")
    
    # 取目标词对应的所有距离值
    target_pos = word2idx[input_word]
    dists = distance_matrix[target_pos]
    
    # 排序取topn,跳过自身
    topn_pos = np.argsort(dists)[1:topn+1]
    return [word_list[pos] for pos in topn_pos]

# 调用示例
print(get_topn_similar_words('movie', word_embeddings, distance_matrix, topn=5))
# 输出示例:['film', 'cinema', 'picture', 'show', 'flick']

注意:如果后续改用余弦相似度计算,只需替换距离矩阵的生成逻辑即可;余弦相似度数值越大代表相似度越高,排序时改为降序,同样跳过自身匹配项即可。


内容的提问来源于stack exchange,提问作者Zwiebak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 15:21:17