You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历字典从pandas DataFrame获取匹配度最高的行

原代码逻辑错误说明

你的代码无法实现需求的核心原因有3点:

  • dataframe[key] == value返回的是对应列所有行的布尔判断序列,不是单个布尔值,直接放在if后会触发pandas的真值判断歧义报错
  • 遍历字典特征时,只要找到第一个有匹配值的列就执行break终止循环,完全没有统计多个特征的匹配情况,无法计算整体匹配度
  • 没有针对每一行统计匹配的特征总个数,自然无法筛选出匹配度最高的结果
实现方案

核心逻辑是跳过name标识字段后,逐行统计和输入字典特征匹配的总数量,匹配数最高的行就是目标结果。

完整实现代码

import pandas as pd

# 初始化示例DataFrame
df = pd.DataFrame(
    [
        ["Apple", "Red", "Heart", "Sweet"],
        ["Banana", "Yellow", "Long", "Sweet"],
        ["Cherry", "Pink", "Circular", "Sour"],
        ["Damson", "Magenta", "Circular", "Sour"],
        ["Eggplant", "Violet", "Long", "Bitter"]
    ],
    columns=["name", "color", "shape", "taste"]
)

new_fruit = {"name": "Tangerine", "color": "Orange", "shape": "Circular", "taste": "Sour"}

# 过滤掉不需要匹配的name字段
match_cols = {k: v for k, v in new_fruit.items() if k != "name"}

# 向量化计算每行的匹配得分(匹配的特征个数)
match_score = sum(df[col] == val for col, val in match_cols.items())

# 筛选得分最高的所有行
result = df[match_score == match_score.max()]
print(result)

运行结果

name    color     shape taste
2   Cherry     Pink  Circular  Sour
3   Damson  Magenta  Circular  Sour

和预期一致,Cherry和Damson两行都匹配了shape、taste两个特征,匹配度最高。
如果数据量很小,也可以用apply写法实现,逻辑更直观:

match_cols = {k: v for k, v in new_fruit.items() if k != "name"}
df["match_score"] = df.apply(lambda row: sum(row[col] == val for col, val in match_cols.items()), axis=1)
result = df[df["match_score"] == df["match_score"].max()].drop(columns="match_score")

内容的提问来源于stack exchange,提问作者Sharp Thwey Thit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.03 05:39:32