计算句子余弦相似度遇ValueError:期望2D数组却得到1D数组
问题:计算句子嵌入余弦相似度时出现ValueError错误
我有11个句子的嵌入向量,形状为(11, 3072),尝试用sklearn的cosine_similarity生成(11,11)的相似度矩阵(每个句子与所有句子的相似度),但运行时出现ValueError: Expected 2D array, got 1D array instead错误。
嵌入向量信息
The shape of embedding (11, 3072) [[-0.02179624 -0.17235152 -0.14017016 ... 0.33180898 0.13701975 -0.2275123 ] [ 0.08176168 0.03396776 -0.00361721 ... -0.06099782 -0.1941497 0.16414282] [ 0.01786027 -0.07074962 0.08268858 ... -0.15433213 0.22098969 -0.05902294] ... [-0.33807683 0.06110802 0.32764304 ... 0.07062552 -0.2734855 -0.01919978] [-0.09536518 0.04956777 0.64503926 ... -0.11085486 -0.36796266 0.2826454 ] [-0.12355942 -0.1552269 -0.01554828 ... -0.14761439 0.17142747 -0.02176587]]
句子示例
document1 = ["sentence a", "sentence b", "sentence c", ...] # 共11个句子
错误代码
# Cosine Similarity from sklearn.metrics.pairwise import cosine_similarity sentences_2d = np.array(document1).reshape(-1,1) similarity_matrix = np.zeros([len(document1), len(document1)]) for i in range(len(sentences_2d)): for j in range(len(sentences_2d)): if i != j: similarity_matrix[i][j] = cosine_similarity(arrcatembed[i], arrcatembed[j])
错误信息
--------------------------------------------------------------------------- ValueError Traceback (most recent call last) <ipython-input-46-e15cce98d633> in <module> 6 for j in range(len(sentences_2d)): 7 if i != j: ----> 8 similarity_matrix[i][j] = cosine_similarity(arrcatembed[i], arrcatembed[j]) 2 frames /usr/local/lib/python3.9/dist-packages/sklearn/utils/validation.py in check_array(array, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, ensure_min_samples, ensure_min_features, estimator, input_name) 900 # If input is 1D raise error 901 if array.ndim == 1: --> 902 raise ValueError( 903 "Expected 2D array, got 1D array instead:\narray={}.\n" 904 "Reshape your data either using array.reshape(-1, 1) if " ValueError: Expected 2D array, got 1D array instead: array=[-0.02179624 -0.17235152 -0.14017016 ... 0.33180898 0.13701975 -0.2275123 ]. Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.
期望生成的相似度矩阵格式
The shape (11, 11) The length 11 [[1. 0.90366799 0.92140669 0.90678644 0.88496917 0.89278495 0.93188739 0.87325549 0.88947386 0.86656564 0.90396279] [0.90366799 1. 0.91544878 0.95543408 0.93818021 0.94250894 0.93432641 0.93418741 0.92931563 0.9156481 0.91719031] [0.92140669 0.91544878 1. 0.92346388 0.91356987 0.93290257 0.94972414 0.90773791 0.92120057 0.90897304 0.92319667] [0.90678644 0.95543408 0.92346388 1. 0.94258463 0.95669407 0.94972783 0.93550926 0.93902498 0.93075407 0.92586052] [0.88496917 0.93818021 0.91356987 0.94258463 1. 0.95144665 0.92863572 0.95595235 0.9522922 0.94791383 0.94201249] [0.89278495 0.94250894 0.93290257 0.95669407 0.95144665 1. 0.95301741 0.95989478 0.95237011 0.94007719 0.93626297] [0.93188739 0.93432641 0.94972414 0.94972783 0.92863572 0.95301741 1. 0.92727625 0.93515086 0.92043686 0.92175251] [0.87325549 0.93418741 0.90773791 0.93550926 0.95595235 0.95989478 0.92727625 1. 0.96572489 0.95371407 0.92973185] [0.88947386 0.92931563 0.92120057 0.93902498 0.9522922 0.95237011 0.93515086 0.96572489 1. 0.95132333 0.9478088 ] [0.86656564 0.9156481 0.90897304 0.93075407 0.94791383 0.94007719 0.92043686 0.95371407 0.95132333 1. 0.92758161] [0.90396279 0.91719031 0.92319667 0.92586052 0.94201249 0.93626297 0.92175251 0.92973185 0.9478088 0.92758161 1. ]]
解决方案
错误原因
cosine_similarity要求输入是2D数组,但代码中arrcatembed[i]取出的是单个1D向量(形状为(3072,)),不符合函数的输入要求。
最优解法:直接传入整个嵌入矩阵
sklearn的cosine_similarity支持直接传入形状为(n, d)的数组,自动计算所有样本间的相似度,返回(n, n)的矩阵,无需手动循环,效率更高:
from sklearn.metrics.pairwise import cosine_similarity import numpy as np # arrcatembed是你的嵌入向量数组,形状(11, 3072) similarity_matrix = cosine_similarity(arrcatembed) # 输出结果 print(f"The shape {similarity_matrix.shape}") print(f"The length {len(similarity_matrix)}") print(similarity_matrix)
循环版本修正(如果必须用循环)
如果一定要保留循环逻辑,需要将每个1D向量转为2D形状(1, 3072):
from sklearn.metrics.pairwise import cosine_similarity import numpy as np similarity_matrix = np.zeros([11, 11]) for i in range(11): for j in range(11): if i == j: # 自身相似度为1 similarity_matrix[i][j] = 1.0 else: # 将1D向量reshape为2D sim_score = cosine_similarity( arrcatembed[i].reshape(1, -1), arrcatembed[j].reshape(1, -1) ) similarity_matrix[i][j] = sim_score[0][0] # 输出结果 print(f"The shape {similarity_matrix.shape}") print(f"The length {len(similarity_matrix)}") print(similarity_matrix)
内容的提问来源于stack exchange,提问作者intodarkmoon
相关产品推荐
相关产品推荐

