You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Fuzzy Wuzzy匹配数据集,如何生成目标DataFrame格式结果?

解决Fuzzy Wuzzy匹配结果整理为DataFrame的问题

假设你的数据结构如下:

  • 原始数据(Raw Data):含待匹配文本列(例如raw_text)的DataFrame
  • 查找表(Lookup Table):含标准文本列(例如standard_text)的DataFrame

步骤1:导入依赖库

import pandas as pd
from fuzzywuzzy import process, fuzz

步骤2:批量匹配并整理结果

不用逐个循环拼接,直接用apply批量处理,将匹配结果转为结构化数据后合并到原始表:

# 提取查找表的标准文本列表
lookup_candidates = lookup_table['standard_text'].tolist()

# 定义匹配函数,返回最优匹配文本和相似度分数
def get_best_match(text):
    # 跳过空值避免报错
    if pd.isna(text) or text.strip() == "":
        return pd.Series([None, None], index=['matched_text', 'match_score'])
    # 提取相似度最高的结果
    match, score = process.extractOne(text, lookup_candidates, scorer=fuzz.token_sort_ratio)
    return pd.Series([match, score], index=['matched_text', 'match_score'])

# 应用函数并合并到原始数据
final_result = raw_data.join(raw_data['raw_text'].apply(get_best_match))

步骤3:扩展多匹配结果(可选)

如果需要返回前N个匹配结果,修改函数即可:

def get_top_n_matches(text, top_n=3):
    if pd.isna(text) or text.strip() == "":
        # 生成空值列
        empty_cols = [None]* (top_n*2)
        cols = [f'matched_text_{i+1}' for i in range(top_n)] + [f'match_score_{i+1}' for i in range(top_n)]
        return pd.Series(empty_cols, index=cols)
    # 提取前N个匹配结果
    matches = process.extract(text, lookup_candidates, scorer=fuzz.token_set_ratio, limit=top_n)
    result = {}
    for idx, (match_text, match_score) in enumerate(matches):
        result[f'matched_text_{idx+1}'] = match_text
        result[f'match_score_{idx+1}'] = match_score
    # 补全不足N个的匹配项为None
    for i in range(len(matches), top_n):
        result[f'matched_text_{i+1}'] = None
        result[f'match_score_{i+1}'] = None
    return pd.Series(result)

# 生成带多匹配结果的表格
multi_match_result = raw_data.join(raw_data['raw_text'].apply(get_top_n_matches))

关键细节

  • 选择合适的匹配器:token_sort_ratio适合语序不同的文本,token_set_ratio适合含重复词的场景,ratio用于整串直接匹配,根据数据特性选择。
  • 空值处理:必须提前判断空文本,避免process.extract抛出异常。

内容的提问来源于stack exchange,提问作者junkxboxx

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.04 12:12:11