如何在Pandas中基于余弦相似度匹配并筛选DataFrame行?
问题描述
现有两个DataFrame:
df1: font_label |font_size | len_words |letter_per_words |text_area_ratio | image_area | Effectiveness | 1 11 7 9.714286 0.046231 310200 | 20.2 2 10.5 8 11 0.0399 310150 19.2 1 11.5 9 10 0.040 310100 21.2
df2: font_label |font_size | len_words |letter_per_words |text_area_ratio | image_area | Effectiveness | 1 12 8 10.5 0.0399 310100 | 21
需要编写函数,传入df2后,从df1中找出与df2余弦相似度最高的行,且该行的Effectiveness值大于df2中的对应值。已导入sklearn.metrics.pairwise.cosine_similarity,预期输出如下:
font_label |font_size | len_words |letter_per_words |text_area_ratio | image_area | Effectiveness | 1 11.5 9 10 0.040 310100 21.2
解决方案
直接使用以下代码实现需求:
import pandas as pd from sklearn.metrics.pairwise import cosine_similarity def find_best_match(df1, df2): # 定义用于计算相似度的特征列(排除Effectiveness) feature_cols = [col for col in df1.columns if col != 'Effectiveness'] # 提取df2的特征向量,调整格式适配cosine_similarity输入要求 df2_features = df2[feature_cols].values.reshape(1, -1) # 计算df1每行与df2的余弦相似度 similarities = cosine_similarity(df1[feature_cols], df2_features).flatten() # 筛选出df1中Effectiveness大于df2对应值的行 df2_eff = df2['Effectiveness'].iloc[0] filtered_rows = df1[df1['Effectiveness'] > df2_eff].copy() # 添加相似度列并排序取最高值的行 filtered_rows['similarity'] = similarities[df1['Effectiveness'] > df2_eff] best_match = filtered_rows.sort_values('similarity', ascending=False).iloc[0].drop('similarity') return best_match.to_frame().T # 测试示例数据 df1 = pd.DataFrame({ 'font_label': [1, 2, 1], 'font_size': [11, 10.5, 11.5], 'len_words': [7, 8, 9], 'letter_per_words': [9.714286, 11, 10], 'text_area_ratio': [0.046231, 0.0399, 0.040], 'image_area': [310200, 310150, 310100], 'Effectiveness': [20.2, 19.2, 21.2] }) df2 = pd.DataFrame({ 'font_label': [1], 'font_size': [12], 'len_words': [8], 'letter_per_words': [10.5], 'text_area_ratio': [0.0399], 'image_area': [310100], 'Effectiveness': [21] }) # 调用函数并输出结果 result = find_best_match(df1, df2) print(result)
代码说明
- 特征列选择:排除
Effectiveness列,仅用其余数值列计算相似度,因为该列是筛选条件而非特征维度。 - 向量格式调整:将df2的特征向量reshape为二维数组,匹配
cosine_similarity的输入格式要求。 - 相似度计算:得到df1每行与df2的相似度数组,转成一维方便后续关联筛选。
- 行筛选:仅保留df1中
Effectiveness值大于df2对应值的行,确保结果符合业务要求。 - 排序取最优:给筛选后的行添加相似度列,按相似度降序排序后取第一行,即为目标结果。
内容的提问来源于stack exchange,提问作者Sasi
相关产品推荐
相关产品推荐

