如何创建函数实现两个DataFrame的国家列值映射及代码优化
代码可读性优化方案
原代码核心问题包括变量命名语义模糊、全局变量依赖风险、特殊别名规则硬编码在分支中维护成本高,优化后代码如下:
import pandas as pd # 统一语义化变量命名,避免无意义的缩写 country_code_df = pd.read_csv('country_codes.csv') registrant_df = pd.read_csv('Registrants.csv') registrant_df.dropna(inplace=True) # 特殊国家别名映射单独抽离为字典,后续新增/修改规则无需改动函数逻辑 COUNTRY_ALIAS_MAP = { "United States": "United States of America", "Korea, South": "South Korea", "Bahamas, The": "The Bahamas" } def get_corrected_country_name(raw_country: str, country_code_df: pd.DataFrame) -> str: # 优先匹配特殊别名 for alias_prefix, standard_name in COUNTRY_ALIAS_MAP.items(): if str(raw_country).startswith(alias_prefix): return standard_name # 精确匹配国家名称库 if country_code_df["COUNTRY_NAME"].str.contains(str(raw_country), na=False).any(): return raw_country # 未匹配到返回标记 return f"not found {raw_country}" # 调用函数生成修正列 registrant_df['country_corrected'] = registrant_df.apply( lambda row: get_corrected_country_name(row['country'], country_code_df), axis=1 ) registrant_df.to_csv('corrected country names.csv', index=False)
精确匹配+模糊匹配逻辑实现
基于编辑距离的模糊匹配推荐使用fuzzywuzzy库实现,先执行安装命令:pip install fuzzywuzzy python-Levenshtein
完整实现逻辑为:先判断原始值是否存在于数据集2的COUNTRY_CODE列,存在直接返回对应国家名;不存在则计算与所有标准国家名的相似度,返回最高匹配结果,可通过阈值过滤低可信度匹配,代码如下:
import pandas as pd from fuzzywuzzy import process # 数据读取 country_code_df = pd.read_csv('country_codes.csv') registrant_df = pd.read_csv('Registrants.csv') registrant_df.dropna(inplace=True) # 特殊别名映射 COUNTRY_ALIAS_MAP = { "United States": "United States of America", "Korea, South": "South Korea", "Bahamas, The": "The Bahamas" } # 相似度阈值,低于该值判定为未匹配,可根据实际场景调整 SIMILARITY_THRESHOLD = 70 # 提前生成映射字典和标准名称列表,避免重复计算提升效率 code_to_country = dict(zip(country_code_df['COUNTRY_CODE'], country_code_df['COUNTRY_NAME'])) standard_country_list = country_code_df['COUNTRY_NAME'].tolist() def get_corrected_country_name(raw_country: str) -> str: raw_country = str(raw_country).strip() # 1. 特殊别名替换 for alias_prefix, standard_name in COUNTRY_ALIAS_MAP.items(): if raw_country.startswith(alias_prefix): return standard_name # 2. 精确匹配国家编码 if raw_country in code_to_country: return code_to_country[raw_country] # 3. 模糊匹配标准国家名 match_result, score = process.extractOne(raw_country, standard_country_list) if score >= SIMILARITY_THRESHOLD: return match_result # 4. 所有匹配失败返回标记 return f"not found {raw_country}" # 生成修正列 registrant_df['country_corrected'] = registrant_df['country'].apply(get_corrected_country_name) registrant_df.to_csv('corrected country names.csv', index=False)
内容的提问来源于stack exchange,提问作者darshika verma
相关产品推荐
相关产品推荐

