如何在Pandas DataFrame中计算句子间相似度?实现方案二及优化方法
计算Pandas DataFrame中句子两两相似度的优化方案
首先给出你的示例数据代码:
import pandas as pd data = [f'Sent {str(i)}' for i in range(10)] df = pd.DataFrame(data=data, columns=['Sentences'])
对应的DataFrame输出:
Sentences 0 Sent 0 1 Sent 1 2 Sent 2 3 Sent 3 4 Sent 4 5 Sent 5 6 Sent 6 7 Sent 7 8 Sent 8 9 Sent 9
方案二的实现方式
方案二的核心是生成所有不重复的两两句子对(即组合数nC2,避免重复计算i-j和j-i),可以借助itertools.combinations实现,步骤如下:
- 导入工具并提取句子列表:
import itertools from sklearn.metrics.pairwise import cosine_similarity from sklearn.feature_extraction.text import TfidfVectorizer sentences = df['Sentences'].tolist()
- 生成组合对并计算相似度:
# 用TF-IDF转换文本为向量(可替换为其他文本表示方法) vectorizer = TfidfVectorizer() tfidf_matrix = vectorizer.fit_transform(sentences) # 生成所有i<j的索引组合,遍历计算相似度 similarity_results = [] for idx1, idx2 in itertools.combinations(range(len(sentences)), 2): sim_score = cosine_similarity(tfidf_matrix[idx1], tfidf_matrix[idx2])[0][0] similarity_results.append({ 'Sentence1': sentences[idx1], 'Sentence2': sentences[idx2], 'Similarity_Score': round(sim_score, 4) }) # 转换为DataFrame查看结果 result_df = pd.DataFrame(similarity_results)
最终的result_df包含所有nC2个不重复的相似度得分,无冗余重复对。
更优的实现方法
如果句子数量较多(如n>1000),循环遍历效率偏低,推荐用向量化计算直接生成相似度矩阵,再提取上三角部分的结果:
# 生成完整相似度矩阵 sim_matrix = cosine_similarity(tfidf_matrix) # 提取上三角矩阵(不含对角线)的索引和得分 import numpy as np upper_triangle = np.triu_indices_from(sim_matrix, k=1) similarity_scores = sim_matrix[upper_triangle] # 整理为结果DataFrame result_df = pd.DataFrame({ 'Sentence1': [sentences[i] for i in upper_triangle[0]], 'Sentence2': [sentences[j] for j in upper_triangle[1]], 'Similarity_Score': [round(score, 4) for score in similarity_scores] })
这种方法利用numpy和sklearn的向量化运算,效率远高于循环,适合大规模数据场景。
注:示例用TF-IDF+余弦相似度,你可根据需求替换为Word2Vec、BERT嵌入等文本表示方法,或Jaccard相似度、编辑距离等算法。
内容的提问来源于stack exchange,提问作者kcats_wolf
相关产品推荐
相关产品推荐

