迭代JSON构建主JSON时遇'Unhashable type: dict'错误的解决方法
合并多结构JSON文件:解决字典不可哈希错误
问题描述
我正在遍历多个JSON字典,将它们的结构复制到一个主JSON字典中,因为被遍历的JSON在长度上略有差异(部分JSON包含更多条目)。
我的JSON文件结构示例:
{ "General info": [ { "section": "General info", "sub-section": "General info", "heading": "Personal Details", "field": "Full Name", "value": "John Doe" }, { "section": "General info", "sub-section": "General info", "heading": "Personal Details", "field": "Address", "value": "69 Doe Lane" } ], "Disclosure": [ { "section": "Disclosure", "sub-section": "Disclosure", "heading": "Agreement", "field": "Signed", "value": "Yes" }, { "section": "Disclosure", "sub-section": "Disclosure", "heading": "Agreement", "field": "Signed2", "value": "Yes" } ], "Assets and Liabilities": [ { "section": "Assets and Liabilities", "sub-section": "Assets", "heading": "Asset1", "field": "Asset value", "value": "£690" }, { "section": "Assets and Liabilities", "sub-section": "Assets", "heading": "Asset1", "field": "Asset type", "value": "Current Account" }, { "section": "Assets and Liabilities", "sub-section": "Assets", "heading": "Asset2", "field": "Asset value", "value": "£42" }, { "section": "Assets and Liabilities", "sub-section": "Assets", "heading": "Asset2", "field": "Asset type", "value": "Stocks and Shares" } ] }
部分JSON文件在Assets and Liabilities键下的heading中包含更多Assets条目,我的目标是遍历所有文件获取最大结构规模后合并。
编写的代码在处理底层字典列表时出现Unhashable type: 'dict'错误,核心代码如下:
from ast import walk import json, os import pandas as pd # first, set the master JSON as the first JSON file in the list of cases master_JSON = r'''C:\Users\SJPJack\documents\JSON UAT\file1.json''' # set the root dir - where all of my JSON files are rootdir = "JSON UAT" # open the master JSON with open(master_JSON) as f: master_JSON = json.load(f) # dump it for inspection with open('master.json', 'w') as f: json.dump(master_JSON, f) # this is just so I can open it in the editor and track changes # define a variable for deleting Client meeting summary from each JSON deleted = 0 # build out our master JSON based on the list of case JSONs # initiate loop for our JSON files for subdir, dirs, file_names in os.walk(rootdir): for fname in file_names: json_path = os.path.join(subdir, fname) data = json.loads(open(json_path).read()) # load current JSON to data variable # iterate for each root key in the JSON for category in list(data.keys()): if category not in list(master_JSON.keys()): master_JSON[category] = data[category] # add missing categories to our master JSON if missing # iterate for each root key in the JSON for category in list(data.keys()): current_posts = data[category] master_posts = master_JSON[category] for current_post in current_posts: for field_name in list(current_post.keys()): current_post['value'] = "" if current_post not in master_posts: master_JSON[current_post] = current_post
解决方案
核心问题
- 直接用
if current_post not in master_posts判断字典是否存在于列表中,而字典是不可哈希类型,无法执行该判断; - 代码最后一行
master_JSON[current_post] = current_post逻辑错误,master_JSON的键是字符串类型,不能用字典作为键。
修复思路
用可哈希的元组作为每个字典条目的唯一标识(由section、sub-section、heading、field四个字段组合而成),通过集合快速判断该结构是否已存在于主JSON中,避免直接操作字典的包含判断。
修复后的完整代码
import json, os # 初始化主JSON为第一个文件 master_json_path = r'C:\Users\SJPJack\documents\JSON UAT\file1.json' with open(master_json_path) as f: master_json = json.load(f) # 保存初始状态用于检查 with open('master_initial.json', 'w') as f: json.dump(master_json, f, indent=2) rootdir = "JSON UAT" # 遍历所有JSON文件 for subdir, dirs, file_names in os.walk(rootdir): for fname in file_names: json_path = os.path.join(subdir, fname) # 跳过初始主文件,避免重复处理 if json_path == master_json_path: continue with open(json_path, 'r') as f: data = json.load(f) # 1. 添加缺失的顶层分类(仅保留结构,值设为空) for category in data.keys(): if category not in master_json: empty_category = [] for item in data[category]: empty_item = item.copy() empty_item['value'] = "" empty_category.append(empty_item) master_json[category] = empty_category # 2. 补充每个分类下缺失的条目结构 for category in data.keys(): if category not in master_json: continue current_items = data[category] master_items = master_json[category] # 提取主JSON中已有条目的唯一标识(元组可哈希) existing_keys = set() for item in master_items: key_tuple = (item['section'], item['sub-section'], item['heading'], item['field']) existing_keys.add(key_tuple) # 遍历当前文件条目,添加缺失的空值结构 for item in current_items: key_tuple = (item['section'], item['sub-section'], item['heading'], item['field']) if key_tuple not in existing_keys: empty_item = item.copy() empty_item['value'] = "" master_items.append(empty_item) existing_keys.add(key_tuple) # 保存最终的主JSON结构 with open('master_final.json', 'w') as f: json.dump(master_json, f, indent=2)
代码说明
- 用元组存储每个条目的唯一标识,解决字典不可哈希的问题;
- 处理顶层分类时,仅复制结构并将
value设为空,避免引入其他文件的实际数据; - 跳过初始主文件,防止重复处理;
- 使用集合存储已有标识,将存在性判断的时间复杂度从O(n)优化为O(1),提升处理效率。
内容的提问来源于stack exchange,提问作者SJPJack
相关产品推荐
相关产品推荐

