You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Pandas计算每行向量的余弦相似度

生成Pandas向量余弦相似度矩阵

你的思路绕了弯路,不需要用pivot_table构造空矩阵再填充,直接提取所有向量计算两两余弦相似度即可,步骤如下:

步骤1:导入必要库

除了pandas,还需要sklearn的余弦相似度工具和numpy处理向量:

import pandas as pd
from sklearn.metrics.pairwise import cosine_similarity
import numpy as np

步骤2:提取向量矩阵

把vec列的列表转换为二维numpy数组,这是余弦相似度计算的标准输入格式:

# 你的原始数据
data = [['apple', 'one', 0.0, [0.047668457, -0.04888916]], ['banana', 'two', 0.0 , [0.0287323, -0.037841797] ], ['qiwi', 'three', 0.0, [0.031051636, -0.05227661]],
        ['orange', 'one', 1.0, [0.0020618439, -0.055389404]], ['mango', 'two', 1.0, [0.0030326843, -0.036193848]], ['strawberry', 'three', 1.0, [0.008613586, -0.06561279]]]
df = pd.DataFrame(data, columns=['word', 'group', 'count', 'vec'])

# 转换为二维数组
vec_matrix = np.array(df['vec'].tolist())

步骤3:计算相似度并构造目标DataFrame

用cosine_similarity直接计算所有向量两两相似度,再用word列的值作为行、列索引:

# 计算余弦相似度
sim_matrix = cosine_similarity(vec_matrix)

# 构造结果DataFrame
result_df = pd.DataFrame(sim_matrix, index=df['word'], columns=df['word']).reset_index()
result_df.rename(columns={'index': 'word'}, inplace=True)

结果示例

运行后得到的result_df与你预期格式一致,前两行如下:

word     apple    banana      qiwi    orange     mango  strawberry
0     apple  1.000000  0.992402  0.972101  0.741462  0.741462    0.800766
1    banana  0.992402  1.000000  0.993574  0.818384  0.844152    0.868376

原代码问题说明

你用pivot_table生成的是对角线上有向量、其余为null的稀疏矩阵,这种结构无法直接传入cosine_similarity——该函数需要完整的向量矩阵作为输入,直接提取所有向量计算才是高效可行的方案。

内容的提问来源于stack exchange,提问作者Rory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 09:30:24