You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中基于余弦相似度匹配并筛选DataFrame行?

问题描述

现有两个DataFrame:

df1:
font_label |font_size | len_words |letter_per_words |text_area_ratio | image_area   | Effectiveness |
    1          11           7          9.714286          0.046231         310200    |    20.2
    2          10.5         8           11               0.0399           310150         19.2
    1          11.5         9           10               0.040            310100         21.2
df2:
font_label |font_size | len_words |letter_per_words |text_area_ratio | image_area   | Effectiveness |
    1          12           8          10.5              0.0399           310100    |    21

需要编写函数,传入df2后,从df1中找出与df2余弦相似度最高的行,且该行的Effectiveness值大于df2中的对应值。已导入sklearn.metrics.pairwise.cosine_similarity,预期输出如下:

font_label |font_size | len_words |letter_per_words |text_area_ratio | image_area   | Effectiveness |
    1          11.5         9           10               0.040            310100         21.2    
解决方案

直接使用以下代码实现需求:

import pandas as pd
from sklearn.metrics.pairwise import cosine_similarity

def find_best_match(df1, df2):
    # 定义用于计算相似度的特征列(排除Effectiveness)
    feature_cols = [col for col in df1.columns if col != 'Effectiveness']
    
    # 提取df2的特征向量,调整格式适配cosine_similarity输入要求
    df2_features = df2[feature_cols].values.reshape(1, -1)
    
    # 计算df1每行与df2的余弦相似度
    similarities = cosine_similarity(df1[feature_cols], df2_features).flatten()
    
    # 筛选出df1中Effectiveness大于df2对应值的行
    df2_eff = df2['Effectiveness'].iloc[0]
    filtered_rows = df1[df1['Effectiveness'] > df2_eff].copy()
    
    # 添加相似度列并排序取最高值的行
    filtered_rows['similarity'] = similarities[df1['Effectiveness'] > df2_eff]
    best_match = filtered_rows.sort_values('similarity', ascending=False).iloc[0].drop('similarity')
    
    return best_match.to_frame().T

# 测试示例数据
df1 = pd.DataFrame({
    'font_label': [1, 2, 1],
    'font_size': [11, 10.5, 11.5],
    'len_words': [7, 8, 9],
    'letter_per_words': [9.714286, 11, 10],
    'text_area_ratio': [0.046231, 0.0399, 0.040],
    'image_area': [310200, 310150, 310100],
    'Effectiveness': [20.2, 19.2, 21.2]
})

df2 = pd.DataFrame({
    'font_label': [1],
    'font_size': [12],
    'len_words': [8],
    'letter_per_words': [10.5],
    'text_area_ratio': [0.0399],
    'image_area': [310100],
    'Effectiveness': [21]
})

# 调用函数并输出结果
result = find_best_match(df1, df2)
print(result)

代码说明

  • 特征列选择:排除Effectiveness列,仅用其余数值列计算相似度,因为该列是筛选条件而非特征维度。
  • 向量格式调整:将df2的特征向量reshape为二维数组,匹配cosine_similarity的输入格式要求。
  • 相似度计算:得到df1每行与df2的相似度数组,转成一维方便后续关联筛选。
  • 行筛选:仅保留df1中Effectiveness值大于df2对应值的行,确保结果符合业务要求。
  • 排序取最优:给筛选后的行添加相似度列,按相似度降序排序后取第一行,即为目标结果。

内容的提问来源于stack exchange,提问作者Sasi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 13:05:24