You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中用for循环按国家分组行政区及KeyError报错解决

解决你的DataFrame分组KeyError问题

嘿,我来帮你搞定这个报错!咱们先拆解问题出在哪:

错误原因分析

  1. KeyError的直接诱因:你用dn['interview_country', 'interview_district']选取多列的方式不对。在pandas里,选取多列必须用双层方括号,单层方括号会把('interview_country', 'interview_district')当成一个单独的列名去查找,而你的DataFrame里显然没有这个列,所以直接抛出KeyError。
  2. 循环遍历逻辑错误:就算列选对了,for d, f in dn[...]的写法也没法正确提取每行的两个字段,得用pandas专门的遍历方法比如iterrows()。
  3. 笔误小坑:你定义了mo_d = [],但最后一个判断分支里写的是zo_d.append(f),这会触发新的NameError,得改成mo_d.append(f)。

修复方案一:改进你的循环写法

先修正列选取方式,再正确遍历每行数据:

sa_d = [] 
ma_d = [] 
za_d = [] 
mo_d = [] 

# 用双层方括号选取多列,再用iterrows遍历每一行(_是占位符,用来接收不需要的行索引)
for _, row in dn[['interview_country', 'interview_district']].iterrows():
    country = row['interview_country']
    district = row['interview_district']
    if country == 'South Africa':
        sa_d.append(district)
    elif country == 'Malawi':
        ma_d.append(district)
    elif country == 'Zambia':
        za_d.append(district)
    elif country == 'Mozambique':
        mo_d.append(district)  # 修正笔误

修复方案二:更高效的Pandas分组方式(强烈推荐)

手动写循环不仅麻烦,数据量大的时候还很慢。用pandas内置的groupby方法可以一键完成分组,代码更简洁高效:

# 按国家分组,将每个国家的行政区聚合为列表,最后转成字典
country_district_map = dn.groupby('interview_country')['interview_district'].agg(list).to_dict()

# 直接通过字典键获取对应国家的行政区列表,第二个参数是默认空列表,避免键不存在时报错
sa_d = country_district_map.get('South Africa', [])
ma_d = country_district_map.get('Malawi', [])
za_d = country_district_map.get('Zambia', [])
mo_d = country_district_map.get('Mozambique', [])

这种方法的优势很明显:

  • 不需要手动维护多个列表,代码更简洁易读
  • Pandas的分组函数是C级别的优化,比Python原生循环快得多
  • 自动适配所有国家,以后新增国家也不用修改代码

最后再提醒下:确认你的DataFrame里确实存在interview_country和interview_district这两个列,没有拼写错误哦!

内容的提问来源于stack exchange,提问作者ckroutz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:12:44