You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于球员姓名合并足球数据表:模糊匹配失效的解决方案求助

解决足球球员长名与昵称的表匹配问题

原代码使用fuzz.ratio全字符比对,无法有效匹配长名(如"Kevin Oghenetega Tamaraebi Bakumo-Abraham")和昵称(如"Tammy Abraham")——两者整体字符重叠度低,导致匹配效果差。以下是针对性改进方案:

1. 替换更适合的模糊匹配打分器

fuzz.ratio是全字符串相似度比对,对长名与昵称这类子串/别名匹配效果不佳,改用以下两种打分器:

  • fuzz.partial_ratio:优先匹配子串,适合昵称是长名一部分的场景
  • fuzz.token_set_ratio:将字符串拆分为单词集合,忽略顺序和重复,能有效匹配包含共同核心词(如姓氏)的名字

修改后的代码:

from fuzzywuzzy import fuzz, process
import pandas as pd

def find_best_match(name, choices):
    # 改用token_set_ratio,兼顾子串和核心词匹配,设置分数阈值过滤低匹配结果
    return process.extractOne(name, choices, scorer=fuzz.token_set_ratio, score_cutoff=70)

# 应用函数匹配
forwards['best_match'] = forwards['Player'].apply(lambda x: find_best_match(x, fifa_fows['long_name']))
forwards['best_name'] = forwards['best_match'].apply(lambda x: x[0] if x else None)

# 合并表并清理中间列
final_forwards = pd.merge(forwards, fifa_fows[['long_name', 'overall']], 
                          left_on='best_name', right_on='long_name', how='left').drop(['best_match', 'best_name'], axis=1)

2. 预处理姓名,提取核心信息

对姓名做标准化处理,减少无关字符/单词的干扰:

  • 统一转小写、去除特殊字符(连字符、点、空格)
  • 提取姓氏和常用名(如取最后一个单词作为姓氏,或提取核心名字片段)

示例预处理代码:

import re

def preprocess_name(name):
    # 转小写,去除非字母字符
    clean_name = re.sub(r'[^a-zA-Z\s]', '', name.lower()).strip()
    words = clean_name.split()
    # 返回姓氏+最后一个名字片段(适配昵称带姓氏的场景)
    return f"{words[-1]} {words[-2]}" if len(words) >=2 else clean_name

# 对两张表的姓名做预处理
fifa_fows['clean_long_name'] = fifa_fows['long_name'].apply(preprocess_name)
forwards['clean_player_name'] = forwards['Player'].apply(preprocess_name)

# 基于预处理后的名字匹配
match_map = dict(zip(fifa_fows['clean_long_name'], fifa_fows['long_name']))
def find_best_match_clean(name, choices):
    match = process.extractOne(name, choices, scorer=fuzz.token_set_ratio, score_cutoff=80)
    return match_map.get(match[0]) if match else None

forwards['best_name'] = forwards['clean_player_name'].apply(lambda x: find_best_match_clean(x, fifa_fows['clean_long_name']))

# 合并表并清理
final_forwards = pd.merge(forwards, fifa_fows[['long_name', 'overall']], 
                          left_on='best_name', right_on='long_name', how='left').drop(['clean_player_name', 'best_name'], axis=1)

3. 构建固定昵称映射表

足球球员的昵称大多固定,可手动整理映射字典,优先用精确匹配,剩余未匹配项再用模糊匹配补充:

# 手动整理常见昵称-长名映射
nickname_map = {
    "Tammy Abraham": "Kevin Oghenetega Tamaraebi Bakumo-Abraham",
    "Mo Salah": "Mohamed Salah Ghaly",
    # 补充更多常见映射
}

# 优先精确匹配
forwards['best_name'] = forwards['Player'].map(nickname_map)

# 对未匹配项用模糊匹配补充
unmatched = forwards[forwards['best_name'].isna()]
unmatched['best_name'] = unmatched['Player'].apply(lambda x: find_best_match(x, fifa_fows['long_name'])[0] if find_best_match(x, fifa_fows['long_name']) else None)

# 更新原表
forwards.loc[unmatched.index, 'best_name'] = unmatched['best_name']

# 最终合并
final_forwards = pd.merge(forwards, fifa_fows[['long_name', 'overall']], 
                          left_on='best_name', right_on='long_name', how='left')

内容的提问来源于stack exchange,提问作者Mocak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 12:56:32