You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

计算句子余弦相似度遇ValueError:期望2D数组却得到1D数组

问题:计算句子嵌入余弦相似度时出现ValueError错误

我有11个句子的嵌入向量,形状为(11, 3072),尝试用sklearn的cosine_similarity生成(11,11)的相似度矩阵(每个句子与所有句子的相似度),但运行时出现ValueError: Expected 2D array, got 1D array instead错误。

嵌入向量信息

The shape of embedding (11, 3072)
[[-0.02179624 -0.17235152 -0.14017016 ...  0.33180898  0.13701975
  -0.2275123 ]
 [ 0.08176168  0.03396776 -0.00361721 ... -0.06099782 -0.1941497
   0.16414282]
 [ 0.01786027 -0.07074962  0.08268858 ... -0.15433213  0.22098969
  -0.05902294]
 ...
 [-0.33807683  0.06110802  0.32764304 ...  0.07062552 -0.2734855
  -0.01919978]
 [-0.09536518  0.04956777  0.64503926 ... -0.11085486 -0.36796266
   0.2826454 ]
 [-0.12355942 -0.1552269  -0.01554828 ... -0.14761439  0.17142747
  -0.02176587]]

句子示例

document1 = ["sentence a", "sentence b", "sentence c", ...] # 共11个句子

错误代码

# Cosine Similarity
from sklearn.metrics.pairwise import cosine_similarity
sentences_2d = np.array(document1).reshape(-1,1)
similarity_matrix = np.zeros([len(document1), len(document1)])
for i in range(len(sentences_2d)):
  for j in range(len(sentences_2d)):
    if i != j:
      similarity_matrix[i][j] = cosine_similarity(arrcatembed[i], arrcatembed[j])

错误信息

---------------------------------------------------------------------------
ValueError                                Traceback (most recent call last)
<ipython-input-46-e15cce98d633> in <module>
      6   for j in range(len(sentences_2d)):
      7     if i != j:
----> 8       similarity_matrix[i][j] = cosine_similarity(arrcatembed[i], arrcatembed[j])

2 frames
/usr/local/lib/python3.9/dist-packages/sklearn/utils/validation.py in check_array(array, accept_sparse, accept_large_sparse, dtype, order, copy, force_all_finite, ensure_2d, allow_nd, ensure_min_samples, ensure_min_features, estimator, input_name)
    900             # If input is 1D raise error
    901             if array.ndim == 1:
--> 902                 raise ValueError(
    903                     "Expected 2D array, got 1D array instead:\narray={}.\n"
    904                     "Reshape your data either using array.reshape(-1, 1) if "

ValueError: Expected 2D array, got 1D array instead:
array=[-0.02179624 -0.17235152 -0.14017016 ...  0.33180898  0.13701975
 -0.2275123 ].
Reshape your data either using array.reshape(-1, 1) if your data has a single feature or array.reshape(1, -1) if it contains a single sample.

期望生成的相似度矩阵格式

The shape (11, 11)
The length 11
[[1.         0.90366799 0.92140669 0.90678644 0.88496917 0.89278495
  0.93188739 0.87325549 0.88947386 0.86656564 0.90396279]
 [0.90366799 1.         0.91544878 0.95543408 0.93818021 0.94250894
  0.93432641 0.93418741 0.92931563 0.9156481  0.91719031]
 [0.92140669 0.91544878 1.         0.92346388 0.91356987 0.93290257
  0.94972414 0.90773791 0.92120057 0.90897304 0.92319667]
 [0.90678644 0.95543408 0.92346388 1.         0.94258463 0.95669407
  0.94972783 0.93550926 0.93902498 0.93075407 0.92586052]
 [0.88496917 0.93818021 0.91356987 0.94258463 1.         0.95144665
  0.92863572 0.95595235 0.9522922  0.94791383 0.94201249]
 [0.89278495 0.94250894 0.93290257 0.95669407 0.95144665 1.
  0.95301741 0.95989478 0.95237011 0.94007719 0.93626297]
 [0.93188739 0.93432641 0.94972414 0.94972783 0.92863572 0.95301741
  1.         0.92727625 0.93515086 0.92043686 0.92175251]
 [0.87325549 0.93418741 0.90773791 0.93550926 0.95595235 0.95989478
  0.92727625 1.         0.96572489 0.95371407 0.92973185]
 [0.88947386 0.92931563 0.92120057 0.93902498 0.9522922  0.95237011
  0.93515086 0.96572489 1.         0.95132333 0.9478088 ]
 [0.86656564 0.9156481  0.90897304 0.93075407 0.94791383 0.94007719
  0.92043686 0.95371407 0.95132333 1.         0.92758161]
 [0.90396279 0.91719031 0.92319667 0.92586052 0.94201249 0.93626297
  0.92175251 0.92973185 0.9478088  0.92758161 1.        ]]

解决方案

错误原因

cosine_similarity要求输入是2D数组,但代码中arrcatembed[i]取出的是单个1D向量(形状为(3072,)),不符合函数的输入要求。

最优解法:直接传入整个嵌入矩阵

sklearn的cosine_similarity支持直接传入形状为(n, d)的数组,自动计算所有样本间的相似度,返回(n, n)的矩阵,无需手动循环,效率更高:

from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

# arrcatembed是你的嵌入向量数组,形状(11, 3072)
similarity_matrix = cosine_similarity(arrcatembed)

# 输出结果
print(f"The shape {similarity_matrix.shape}")
print(f"The length {len(similarity_matrix)}")
print(similarity_matrix)

循环版本修正(如果必须用循环)

如果一定要保留循环逻辑,需要将每个1D向量转为2D形状(1, 3072):

from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

similarity_matrix = np.zeros([11, 11])
for i in range(11):
    for j in range(11):
        if i == j:
            # 自身相似度为1
            similarity_matrix[i][j] = 1.0
        else:
            # 将1D向量reshape为2D
            sim_score = cosine_similarity(
                arrcatembed[i].reshape(1, -1), 
                arrcatembed[j].reshape(1, -1)
            )
            similarity_matrix[i][j] = sim_score[0][0]

# 输出结果
print(f"The shape {similarity_matrix.shape}")
print(f"The length {len(similarity_matrix)}")
print(similarity_matrix)

内容的提问来源于stack exchange,提问作者intodarkmoon

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 06:12:03