如何将CSV数据加载为嵌套字典以快速查询模板名称?
问题描述
给定如下结构的CSV文件:
Category, Sub-Category, Template Name Health Check,CPU,checkCPU Health Check,Memory,checkMemory Service Request,Reboot Device, rebootDevice Service Request,Check CPU,checkCPU-SR
需要编写Python脚本,实现通过Category和Sub-Category快速获取对应的Template Name(数据会随时间更新)。当前已能通过读取CSV后循环查询实现,但希望将CSV转换为{Category: {Sub-Category: Template Name}}格式的嵌套字典,直接像操作JSON那样取值,尝试过csv.DictReader和Pandas但未成功,寻求可行方案。
解决方案
方法一:使用标准库csv手动构建嵌套字典
无需额外依赖,用Python内置csv模块读取数据,逐步构建嵌套结构:
import csv def csv_to_nested_dict(csv_file_path): nested_dict = {} with open(csv_file_path, mode='r', newline='', encoding='utf-8') as file: reader = csv.DictReader(file) for row in reader: # 去除字段名和值的首尾空格(适配CSV中带空格的列名/值) category = row['Category'].strip() sub_category = row[' Sub-Category'].strip() template_name = row[' Template Name'].strip() # 初始化第一层分类字典 if category not in nested_dict: nested_dict[category] = {} # 填充第二层子分类对应的模板名 nested_dict[category][sub_category] = template_name return nested_dict # 使用示例 template_map = csv_to_nested_dict('templates.csv') print(template_map['Health Check']['CPU']) # 输出: checkCPU print(template_map['Service Request']['Reboot Device']) # 输出: rebootDevice
方法二:使用Pandas构建嵌套字典
若已依赖Pandas,可通过索引设置和字典转换快速实现:
import pandas as pd def pandas_to_nested_dict(csv_file_path): df = pd.read_csv(csv_file_path) # 批量去除所有字符串类型字段的首尾空格 df = df.apply(lambda x: x.str.strip() if x.dtype == "object" else x) # 生成双层索引的字典,再转换为嵌套结构 raw_dict = df.set_index(['Category', 'Sub-Category'])['Template Name'].to_dict() nested_dict = {} for (cat, sub_cat), name in raw_dict.items(): if cat not in nested_dict: nested_dict[cat] = {} nested_dict[cat][sub_cat] = name return nested_dict # 使用示例 template_map = pandas_to_nested_dict('templates.csv') print(template_map['Health Check']['Memory']) # 输出: checkMemory
关键注意点
- CSV中部分列名(如
Sub-Category)和值带有首尾空格,必须通过strip()处理,否则会出现键不匹配的问题。 - 若存在同一Category+Sub-Category重复的行,后读取的数据会覆盖之前的值,需根据业务需求提前处理重复数据。
内容的提问来源于stack exchange,提问作者ChadDa3mon
相关产品推荐
相关产品推荐

