You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

迭代JSON构建主JSON时遇'Unhashable type: dict'错误的解决方法

合并多结构JSON文件:解决字典不可哈希错误

问题描述

我正在遍历多个JSON字典,将它们的结构复制到一个主JSON字典中,因为被遍历的JSON在长度上略有差异(部分JSON包含更多条目)。

我的JSON文件结构示例:

{
  "General info": [
    {
      "section": "General info",
      "sub-section": "General info",
      "heading": "Personal Details",
      "field": "Full Name",
      "value": "John Doe"
    },
    {
      "section": "General info",
      "sub-section": "General info",
      "heading": "Personal Details",
      "field": "Address",
      "value": "69 Doe Lane"
    }
  ],
  "Disclosure": [
    {
      "section": "Disclosure",
      "sub-section": "Disclosure",
      "heading": "Agreement",
      "field": "Signed",
      "value": "Yes"
    },
    {
      "section": "Disclosure",
      "sub-section": "Disclosure",
      "heading": "Agreement",
      "field": "Signed2",
      "value": "Yes"
    }
  ],
  "Assets and Liabilities": [
    {
      "section": "Assets and Liabilities",
      "sub-section": "Assets",
      "heading": "Asset1",
      "field": "Asset value",
      "value": "£690"
    },
    {
      "section": "Assets and Liabilities",
      "sub-section": "Assets",
      "heading": "Asset1",
      "field": "Asset type",
      "value": "Current Account"
    },
    {
      "section": "Assets and Liabilities",
      "sub-section": "Assets",
      "heading": "Asset2",
      "field": "Asset value",
      "value": "£42"
    },
    {
      "section": "Assets and Liabilities",
      "sub-section": "Assets",
      "heading": "Asset2",
      "field": "Asset type",
      "value": "Stocks and Shares"
    }
  ]
}

部分JSON文件在Assets and Liabilities键下的heading中包含更多Assets条目,我的目标是遍历所有文件获取最大结构规模后合并。

编写的代码在处理底层字典列表时出现Unhashable type: 'dict'错误,核心代码如下:

from ast import walk
import json, os
import pandas as pd

# first, set the master JSON as the first JSON file in the list of cases
master_JSON = r'''C:\Users\SJPJack\documents\JSON UAT\file1.json'''

# set the root dir - where all of my JSON files are
rootdir = "JSON UAT"

# open the master JSON
with open(master_JSON) as f:
    master_JSON = json.load(f)

# dump it for inspection
with open('master.json', 'w') as f:
    json.dump(master_JSON, f) # this is just so I can open it in the editor and track changes

# define a variable for deleting Client meeting summary from each JSON
deleted = 0

# build out our master JSON based on the list of case JSONs
# initiate loop for our JSON files
for subdir, dirs, file_names in os.walk(rootdir):
    for fname in file_names:
        json_path = os.path.join(subdir, fname)
        data = json.loads(open(json_path).read()) # load current JSON to data variable

        # iterate for each root key in the JSON 
        for category in list(data.keys()):
            if category not in list(master_JSON.keys()):
                master_JSON[category] = data[category] # add missing categories to our master JSON if missing

        # iterate for each root key in the JSON 
        for category in list(data.keys()):
            current_posts = data[category]
            master_posts = master_JSON[category]

            for current_post in current_posts:
                for field_name in list(current_post.keys()):
                    current_post['value'] = ""
                
                if current_post not in master_posts:
                    master_JSON[current_post] = current_post

解决方案

核心问题

  1. 直接用if current_post not in master_posts判断字典是否存在于列表中,而字典是不可哈希类型,无法执行该判断;
  2. 代码最后一行master_JSON[current_post] = current_post逻辑错误,master_JSON的键是字符串类型,不能用字典作为键。

修复思路

用可哈希的元组作为每个字典条目的唯一标识(由section、sub-section、heading、field四个字段组合而成),通过集合快速判断该结构是否已存在于主JSON中,避免直接操作字典的包含判断。

修复后的完整代码

import json, os

# 初始化主JSON为第一个文件
master_json_path = r'C:\Users\SJPJack\documents\JSON UAT\file1.json'
with open(master_json_path) as f:
    master_json = json.load(f)

# 保存初始状态用于检查
with open('master_initial.json', 'w') as f:
    json.dump(master_json, f, indent=2)

rootdir = "JSON UAT"

# 遍历所有JSON文件
for subdir, dirs, file_names in os.walk(rootdir):
    for fname in file_names:
        json_path = os.path.join(subdir, fname)
        # 跳过初始主文件,避免重复处理
        if json_path == master_json_path:
            continue
            
        with open(json_path, 'r') as f:
            data = json.load(f)

        # 1. 添加缺失的顶层分类(仅保留结构,值设为空)
        for category in data.keys():
            if category not in master_json:
                empty_category = []
                for item in data[category]:
                    empty_item = item.copy()
                    empty_item['value'] = ""
                    empty_category.append(empty_item)
                master_json[category] = empty_category

        # 2. 补充每个分类下缺失的条目结构
        for category in data.keys():
            if category not in master_json:
                continue
                
            current_items = data[category]
            master_items = master_json[category]
            
            # 提取主JSON中已有条目的唯一标识(元组可哈希)
            existing_keys = set()
            for item in master_items:
                key_tuple = (item['section'], item['sub-section'], item['heading'], item['field'])
                existing_keys.add(key_tuple)
                
            # 遍历当前文件条目,添加缺失的空值结构
            for item in current_items:
                key_tuple = (item['section'], item['sub-section'], item['heading'], item['field'])
                if key_tuple not in existing_keys:
                    empty_item = item.copy()
                    empty_item['value'] = ""
                    master_items.append(empty_item)
                    existing_keys.add(key_tuple)

# 保存最终的主JSON结构
with open('master_final.json', 'w') as f:
    json.dump(master_json, f, indent=2)

代码说明

  • 用元组存储每个条目的唯一标识,解决字典不可哈希的问题;
  • 处理顶层分类时,仅复制结构并将value设为空,避免引入其他文件的实际数据;
  • 跳过初始主文件,防止重复处理;
  • 使用集合存储已有标识,将存在性判断的时间复杂度从O(n)优化为O(1),提升处理效率。

内容的提问来源于stack exchange,提问作者SJPJack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 14:40:20