Python读取3列CSV构建嵌套字典遇键覆盖问题求助
解决嵌套字典覆盖问题
你的代码核心问题是每次循环都会直接用新字典完全替换同一Region对应的旧字典,导致每个Region最终只保留最后一条数据。下面提供几种可行的解决方案:
方案一:检查Region是否存在再更新
先判断当前Region是否已在data_dict中,不存在则初始化空字典,存在则直接向该字典添加键值对:
with open(self.file_name, encoding='iso-8859-1') as file: reader = csv.DictReader(file) for line in reader: region = line['Region'] age_range = line['Range'] population = line['Population'] if region not in self.data_dict: self.data_dict[region] = {} self.data_dict[region][age_range] = population
方案二:使用dict.setdefault()简化代码
setdefault()方法会自动在键不存在时设置默认值(这里是空字典),存在则返回已有值,一行代码完成判断与更新:
with open(self.file_name, encoding='iso-8859-1') as file: reader = csv.DictReader(file) for line in reader: region = line['Region'] age_range = line['Range'] population = line['Population'] self.data_dict.setdefault(region, {})[age_range] = population
方案三:使用collections.defaultdict
提前将data_dict初始化为defaultdict,指定默认值为字典类型,无需手动判断Region是否存在:
from collections import defaultdict # 初始化data_dict,默认值为空字典 self.data_dict = defaultdict(dict) with open(self.file_name, encoding='iso-8859-1') as file: reader = csv.DictReader(file) for line in reader: region = line['Region'] age_range = line['Range'] population = line['Population'] self.data_dict[region][age_range] = population
以上三种方案均可实现需求:每个Region下保存所有Range对应的Population,不会出现覆盖问题。处理示例数据后,self.data_dict["Region1"]["23"]会返回"1299",self.data_dict["Region2"]["45"]返回"3008"。
内容的提问来源于stack exchange,提问作者andsemenov
相关产品推荐
相关产品推荐

