You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何解决DataFrame模糊匹配全量关联问题,实现行内匹配并生成多索引结果

行对行模糊匹配DataFrame字符串与元组元素的实现方案

问题背景

需要对特定结构的DataFrame执行行内模糊匹配:

  • fruits列:每行仅包含单个字符串
  • choices列:每行是包含多个字符串的元组
    要求每行的fruits元素仅与该行choices元组内的每个元素分别计算模糊匹配度,而非生成所有fruits与所有choices元组的交叉匹配结果。需要计算partial_ratio、token_sort_ratio、token_set_ratio三类匹配值,并生成多索引格式的输出。

原始DataFrame定义:

import pandas as pd

df = pd.DataFrame(
    {
        "id": [1, 2, 3, 4, 5, 6],
        "fruits": ["apple", "apples", "orange", "apple tree", "oranges", "mango"],
        "choices": [
            ("app", "apull", "apple"),
            ("app", "apull", "apple", "appple"),
            ("orange", "org"),
            ("apple",),  # 注意:单个元素的元组必须加逗号,否则会被识别为字符串
            ("oranges", "orang"),
            ("mango",),
        ],
    }
)

此前使用pd.MultiIndex.from_product生成了全交叉匹配结果,导致数据量过大,不符合需求:

# 错误实现:生成所有fruits与choices的交叉组合
compare = pd.MultiIndex.from_product([df['fruits'], df['choices']]).to_series()

解决方案

使用apply逐行处理,对每行的fruits和choices元组内元素逐一计算匹配值,再合并结果生成目标格式:

完整代码

import pandas as pd
from fuzzywuzzy import fuzz

# 修正原DataFrame中单个元素的元组格式(避免被识别为字符串)
df['choices'] = df['choices'].apply(lambda x: (x,) if isinstance(x, str) else x)

# 定义行处理函数,计算当前行所有匹配值
def process_single_row(row):
    target_fruit = row['fruits']
    choice_list = row['choices']
    match_results = []
    
    for choice in choice_list:
        match_results.append({
            'id': row['id'],
            'fruits': target_fruit,
            'choice': choice,
            'partial_ratio': fuzz.partial_ratio(target_fruit, choice),
            'token_sort_ratio': fuzz.token_sort_ratio(target_fruit, choice),
            'token_set_ratio': fuzz.token_set_ratio(target_fruit, choice)
        })
    
    return pd.DataFrame(match_results)

# 逐行处理并合并所有结果
final_result = pd.concat(df.apply(process_single_row, axis=1).tolist(), ignore_index=True)

# 可选:设置多索引以匹配期望格式
final_result = final_result.set_index(['id', 'fruits', 'choice'])

# 打印结果
print(final_result)

输出示例

partial_ratio  token_sort_ratio  token_set_ratio
id fruits     choice                                                       
1  apple      app                    80                67                67
              apull                  60                57                57
              apple                 100               100               100
2  apples     app                    75                67                67
              apull                  50                50                50
              apple                  90                91                91
              appple                 80                82                82
3  orange     orange               100               100               100
              org                    67                67                67
4  apple tree apple                 71                71                71
5  oranges    oranges              100               100               100
              orang                  89                89                89
6  mango      mango                100               100               100

内容的提问来源于stack exchange,提问作者lordgriffith

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.11 14:20:57