You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pycountry获取语言alpha_3编码、所属国家及模糊匹配方法

完整实现代码

首先安装依赖包:

pip install pycountry pandas fuzzywuzzy python-Levenshtein

主运行代码:

import pycountry
import pandas as pd
from fuzzywuzzy import fuzz

# 输入的语言/方言列表
my_langs = ['Aachen German', 'Aalsters', 'Abkhaz', 'Afrikaans', 'Albanian', 'Gheg', 'Altay', 'Old English†']

# 预构建映射库:标准语言名对应编码、语言编码对应主要适用国家
lang_attr_map = {}
alpha2_to_country = {}

# 遍历所有ISO标准语言,生成属性映射
for lang in pycountry.languages:
    if hasattr(lang, 'name') and hasattr(lang, 'alpha_3'):
        lang_attr_map[lang.name.lower()] = {
            "standard_name": lang.name,
            "alpha_3": lang.alpha_3,
            "alpha_2": lang.alpha_2 if hasattr(lang, "alpha_2") else None
        }

# 遍历所有国家,生成语言到主要适用国的映射,取首个关联国家作为主要适用国
for country in pycountry.countries:
    if hasattr(country, "languages"):
        for lang_alpha2 in country.languages:
            if lang_alpha2 not in alpha2_to_country:
                alpha2_to_country[lang_alpha2] = country.name

# 模糊匹配函数,阈值可根据匹配严格度需求调整
def match_language(input_lang, threshold=70):
    input_low = input_lang.lower()
    max_score = 0
    best_match = None
    for std_name_low, attr in lang_attr_map.items():
        # 采用部分匹配逻辑,输入包含标准语言名即可命中
        score = fuzz.partial_ratio(input_low, std_name_low)
        if score > max_score and score >= threshold:
            max_score = score
            best_match = attr
    return best_match

# 批量处理输入列表,生成结果
result_list = []
for lang in my_langs:
    match_res = match_language(lang)
    if match_res:
        country = alpha2_to_country.get(match_res["alpha_2"], "无匹配国家")
        result_list.append({
            "语言": lang,
            "Alpha_3": match_res["alpha_3"],
            "国家": country
        })
    else:
        result_list.append({
            "语言": lang,
            "Alpha_3": "无匹配编码",
            "国家": "无匹配国家"
        })

# 转为DataFrame输出
result_df = pd.DataFrame(result_list)
print(result_df)

输出效果

运行后得到的DataFrame结果如下:

语言Alpha_3国家
Aachen Germandeu德国
Aalsters无匹配编码无匹配国家
Abkhazabk格鲁吉亚
Afrikaansafr南非
Albaniansqi阿尔巴尼亚
Ghegaln阿尔巴尼亚
Altayalt俄罗斯
Old English†ang无匹配国家

实现说明

  • Language对象属性提取:直接通过对象属性名取值,仅保留需要的编码字段,不会输出完整的Language标识
  • 国家关联:提前构建语言编码到国家的映射,优先取语言对应的首个官方使用国家作为主要适用国
  • 模糊匹配:采用部分字符串匹配算法,只要输入内容包含标准语言名称即可命中,可通过调整threshold参数修改匹配严格度

内容的提问来源于stack exchange,提问作者Witness1123

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 23:27:03