如何从pandas DataFrame中提取country_code与country列的映射字典
Pandas提取两列映射字典的实现方法
你可以用以下两种方法实现需求,优先推荐第一种更稳妥:
方法1:去重后转索引生成字典(推荐)
该方法不受列顺序、数据异常波动影响,结果可靠性更高:
d = df.drop_duplicates(subset=['country_code'])[['country_code', 'country']].set_index('country_code')['country'].to_dict()
执行后得到的d就是你预期的输出结构:{'arg':'argentina', 'bra':'brazil', 'eng':'england'}
方法2:zip打包去重结果
如果你确认数据中country_code和country是严格一一对应的关系,可以用更简洁的写法:
d = dict(zip(df['country_code'].unique(), df['country'].unique()))
注意:如果存在同编码对应不同国家的异常数据,该方法会出现映射错误,数据质量不确定的情况下不建议使用。
内容的提问来源于stack exchange,提问作者Danish
相关产品推荐
相关产品推荐

