Pandas条件赋值异常:疫苗数据大陆分类错误/空白问题求助
问题排查与解决方案
核心原因分析
你的if elif逻辑出现分类错误、数据留空(比如ZWE未识别),大概率是以下几个问题:
- 硬编码的国家代码列表不全:ZWE(津巴布韦)没被加入非洲的判断列表中
- 大小写不匹配:CSV中的国家代码是小写(如
zwe),但代码中只判断了大写(如ZWE) - 逻辑顺序问题:某些国家代码被提前的条件错误匹配到其他大洲
- 边缘国家/地区未处理:部分国家/地区的代码不在判断范围内,导致留空
修复方案
方案1:用字典映射替代冗长的if elif(推荐)
字典映射比if elif更易维护,且能避免顺序问题,同时统一处理大小写:
# 预定义完整的国家代码-大洲映射(补全你需要的所有国家代码) continent_map = { # 非洲包含ZWE(津巴布韦) 'Africa': ['ZWE', 'EGY', 'NGA', 'KEN', 'ETH', 'SAU', 'MAR', ...], 'North America': ['USA', 'CAN', 'MEX', 'CUB', ...], 'South America': ['BRA', 'ARG', 'CHL', 'PER', ...], 'Asia': ['CHN', 'IND', 'JPN', 'KOR', ...], 'Europe': ['DEU', 'FRA', 'GBR', 'ITA', ...], 'Oceania': ['AUS', 'NZL', 'PNG', ...] } # 转换为扁平化的小写映射,支持不区分大小写匹配 flat_continent_map = { code.lower(): continent for continent, codes in continent_map.items() for code in codes } def assign_continent(country_code): # 去除首尾空格并统一转为小写 cleaned_code = country_code.strip().lower() # 返回对应大洲,未匹配则标记为Unknown return flat_continent_map.get(cleaned_code, 'Unknown') # 测试ZWE print(assign_continent('ZWE')) # 输出:Africa print(assign_continent('zwe')) # 输出:Africa
方案2:用专业库自动匹配(更精准)
如果不想手动维护国家代码列表,可以用pycountry和pycountry-convert库自动完成大洲匹配:
- 先安装依赖:
pip install pycountry pycountry-convert
- 实现代码:
import pycountry from pycountry_convert import country_alpha2_to_continent_code, country_alpha3_to_country_alpha2 def get_continent(country_code): try: # 处理三位字母代码,转为两位标准代码 if len(country_code) == 3: alpha2 = country_alpha3_to_country_alpha2(country_code.strip().upper()) else: alpha2 = country_code.strip().upper() # 转换为大洲代码,再映射为大洲名称 continent_code = country_alpha2_to_continent_code(alpha2) continent_name = { 'AF': 'Africa', 'AS': 'Asia', 'EU': 'Europe', 'NA': 'North America', 'OC': 'Oceania', 'SA': 'South America' }[continent_code] return continent_name except Exception as e: # 打印未识别的代码便于排查 print(f"无法识别国家代码: {country_code}, 错误: {str(e)}") return 'Unknown' # 测试ZWE print(get_continent('ZWE')) # 输出:Africa
额外排查步骤
- 检查CSV中的国家代码:用
print(df['country_code'].unique())输出所有唯一代码,确认是否有拼写错误(比如ZWE写成ZWD) - 验证原
if elif逻辑:检查是否有ZWE的判断,以及是否被其他条件提前匹配 - 处理空值:CSV中可能存在空的国家代码,需提前用
df['country_code'] = df['country_code'].fillna('Unknown')处理
内容的提问来源于stack exchange,提问作者Furkan Ayhan Gülekoğlu
相关产品推荐
相关产品推荐

