基于球员姓名合并足球数据表:模糊匹配失效的解决方案求助
解决足球球员长名与昵称的表匹配问题
原代码使用fuzz.ratio全字符比对,无法有效匹配长名(如"Kevin Oghenetega Tamaraebi Bakumo-Abraham")和昵称(如"Tammy Abraham")——两者整体字符重叠度低,导致匹配效果差。以下是针对性改进方案:
1. 替换更适合的模糊匹配打分器
fuzz.ratio是全字符串相似度比对,对长名与昵称这类子串/别名匹配效果不佳,改用以下两种打分器:
fuzz.partial_ratio:优先匹配子串,适合昵称是长名一部分的场景fuzz.token_set_ratio:将字符串拆分为单词集合,忽略顺序和重复,能有效匹配包含共同核心词(如姓氏)的名字
修改后的代码:
from fuzzywuzzy import fuzz, process import pandas as pd def find_best_match(name, choices): # 改用token_set_ratio,兼顾子串和核心词匹配,设置分数阈值过滤低匹配结果 return process.extractOne(name, choices, scorer=fuzz.token_set_ratio, score_cutoff=70) # 应用函数匹配 forwards['best_match'] = forwards['Player'].apply(lambda x: find_best_match(x, fifa_fows['long_name'])) forwards['best_name'] = forwards['best_match'].apply(lambda x: x[0] if x else None) # 合并表并清理中间列 final_forwards = pd.merge(forwards, fifa_fows[['long_name', 'overall']], left_on='best_name', right_on='long_name', how='left').drop(['best_match', 'best_name'], axis=1)
2. 预处理姓名,提取核心信息
对姓名做标准化处理,减少无关字符/单词的干扰:
- 统一转小写、去除特殊字符(连字符、点、空格)
- 提取姓氏和常用名(如取最后一个单词作为姓氏,或提取核心名字片段)
示例预处理代码:
import re def preprocess_name(name): # 转小写,去除非字母字符 clean_name = re.sub(r'[^a-zA-Z\s]', '', name.lower()).strip() words = clean_name.split() # 返回姓氏+最后一个名字片段(适配昵称带姓氏的场景) return f"{words[-1]} {words[-2]}" if len(words) >=2 else clean_name # 对两张表的姓名做预处理 fifa_fows['clean_long_name'] = fifa_fows['long_name'].apply(preprocess_name) forwards['clean_player_name'] = forwards['Player'].apply(preprocess_name) # 基于预处理后的名字匹配 match_map = dict(zip(fifa_fows['clean_long_name'], fifa_fows['long_name'])) def find_best_match_clean(name, choices): match = process.extractOne(name, choices, scorer=fuzz.token_set_ratio, score_cutoff=80) return match_map.get(match[0]) if match else None forwards['best_name'] = forwards['clean_player_name'].apply(lambda x: find_best_match_clean(x, fifa_fows['clean_long_name'])) # 合并表并清理 final_forwards = pd.merge(forwards, fifa_fows[['long_name', 'overall']], left_on='best_name', right_on='long_name', how='left').drop(['clean_player_name', 'best_name'], axis=1)
3. 构建固定昵称映射表
足球球员的昵称大多固定,可手动整理映射字典,优先用精确匹配,剩余未匹配项再用模糊匹配补充:
# 手动整理常见昵称-长名映射 nickname_map = { "Tammy Abraham": "Kevin Oghenetega Tamaraebi Bakumo-Abraham", "Mo Salah": "Mohamed Salah Ghaly", # 补充更多常见映射 } # 优先精确匹配 forwards['best_name'] = forwards['Player'].map(nickname_map) # 对未匹配项用模糊匹配补充 unmatched = forwards[forwards['best_name'].isna()] unmatched['best_name'] = unmatched['Player'].apply(lambda x: find_best_match(x, fifa_fows['long_name'])[0] if find_best_match(x, fifa_fows['long_name']) else None) # 更新原表 forwards.loc[unmatched.index, 'best_name'] = unmatched['best_name'] # 最终合并 final_forwards = pd.merge(forwards, fifa_fows[['long_name', 'overall']], left_on='best_name', right_on='long_name', how='left')
内容的提问来源于stack exchange,提问作者Mocak
相关产品推荐
相关产品推荐

