You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何创建函数实现两个DataFrame的国家列值映射及代码优化

代码可读性优化方案

原代码核心问题包括变量命名语义模糊、全局变量依赖风险、特殊别名规则硬编码在分支中维护成本高,优化后代码如下:

import pandas as pd

# 统一语义化变量命名,避免无意义的缩写
country_code_df = pd.read_csv('country_codes.csv')
registrant_df = pd.read_csv('Registrants.csv')
registrant_df.dropna(inplace=True)

# 特殊国家别名映射单独抽离为字典,后续新增/修改规则无需改动函数逻辑
COUNTRY_ALIAS_MAP = {
    "United States": "United States of America",
    "Korea, South": "South Korea",
    "Bahamas, The": "The Bahamas"
}

def get_corrected_country_name(raw_country: str, country_code_df: pd.DataFrame) -> str:
    # 优先匹配特殊别名
    for alias_prefix, standard_name in COUNTRY_ALIAS_MAP.items():
        if str(raw_country).startswith(alias_prefix):
            return standard_name
    # 精确匹配国家名称库
    if country_code_df["COUNTRY_NAME"].str.contains(str(raw_country), na=False).any():
        return raw_country
    # 未匹配到返回标记
    return f"not found {raw_country}"

# 调用函数生成修正列
registrant_df['country_corrected'] = registrant_df.apply(
    lambda row: get_corrected_country_name(row['country'], country_code_df), 
    axis=1
)
registrant_df.to_csv('corrected country names.csv', index=False)
精确匹配+模糊匹配逻辑实现

基于编辑距离的模糊匹配推荐使用fuzzywuzzy库实现,先执行安装命令:
pip install fuzzywuzzy python-Levenshtein

完整实现逻辑为:先判断原始值是否存在于数据集2的COUNTRY_CODE列,存在直接返回对应国家名;不存在则计算与所有标准国家名的相似度,返回最高匹配结果,可通过阈值过滤低可信度匹配,代码如下:

import pandas as pd
from fuzzywuzzy import process

# 数据读取
country_code_df = pd.read_csv('country_codes.csv')
registrant_df = pd.read_csv('Registrants.csv')
registrant_df.dropna(inplace=True)

# 特殊别名映射
COUNTRY_ALIAS_MAP = {
    "United States": "United States of America",
    "Korea, South": "South Korea",
    "Bahamas, The": "The Bahamas"
}
# 相似度阈值,低于该值判定为未匹配,可根据实际场景调整
SIMILARITY_THRESHOLD = 70
# 提前生成映射字典和标准名称列表,避免重复计算提升效率
code_to_country = dict(zip(country_code_df['COUNTRY_CODE'], country_code_df['COUNTRY_NAME']))
standard_country_list = country_code_df['COUNTRY_NAME'].tolist()

def get_corrected_country_name(raw_country: str) -> str:
    raw_country = str(raw_country).strip()
    # 1. 特殊别名替换
    for alias_prefix, standard_name in COUNTRY_ALIAS_MAP.items():
        if raw_country.startswith(alias_prefix):
            return standard_name
    # 2. 精确匹配国家编码
    if raw_country in code_to_country:
        return code_to_country[raw_country]
    # 3. 模糊匹配标准国家名
    match_result, score = process.extractOne(raw_country, standard_country_list)
    if score >= SIMILARITY_THRESHOLD:
        return match_result
    # 4. 所有匹配失败返回标记
    return f"not found {raw_country}"

# 生成修正列
registrant_df['country_corrected'] = registrant_df['country'].apply(get_corrected_country_name)
registrant_df.to_csv('corrected country names.csv', index=False)

内容的提问来源于stack exchange,提问作者darshika verma

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.28 11:06:03