如何使用Pandas计算每行向量的余弦相似度
生成Pandas向量余弦相似度矩阵
你的思路绕了弯路,不需要用pivot_table构造空矩阵再填充,直接提取所有向量计算两两余弦相似度即可,步骤如下:
步骤1:导入必要库
除了pandas,还需要sklearn的余弦相似度工具和numpy处理向量:
import pandas as pd from sklearn.metrics.pairwise import cosine_similarity import numpy as np
步骤2:提取向量矩阵
把vec列的列表转换为二维numpy数组,这是余弦相似度计算的标准输入格式:
# 你的原始数据 data = [['apple', 'one', 0.0, [0.047668457, -0.04888916]], ['banana', 'two', 0.0 , [0.0287323, -0.037841797] ], ['qiwi', 'three', 0.0, [0.031051636, -0.05227661]], ['orange', 'one', 1.0, [0.0020618439, -0.055389404]], ['mango', 'two', 1.0, [0.0030326843, -0.036193848]], ['strawberry', 'three', 1.0, [0.008613586, -0.06561279]]] df = pd.DataFrame(data, columns=['word', 'group', 'count', 'vec']) # 转换为二维数组 vec_matrix = np.array(df['vec'].tolist())
步骤3:计算相似度并构造目标DataFrame
用cosine_similarity直接计算所有向量两两相似度,再用word列的值作为行、列索引:
# 计算余弦相似度 sim_matrix = cosine_similarity(vec_matrix) # 构造结果DataFrame result_df = pd.DataFrame(sim_matrix, index=df['word'], columns=df['word']).reset_index() result_df.rename(columns={'index': 'word'}, inplace=True)
结果示例
运行后得到的result_df与你预期格式一致,前两行如下:
word apple banana qiwi orange mango strawberry 0 apple 1.000000 0.992402 0.972101 0.741462 0.741462 0.800766 1 banana 0.992402 1.000000 0.993574 0.818384 0.844152 0.868376
原代码问题说明
你用pivot_table生成的是对角线上有向量、其余为null的稀疏矩阵,这种结构无法直接传入cosine_similarity——该函数需要完整的向量矩阵作为输入,直接提取所有向量计算才是高效可行的方案。
内容的提问来源于stack exchange,提问作者Rory
相关产品推荐
相关产品推荐

