如何在Python中打印嵌套JSON的全部重复键(含完整路径及可选值)
处理JSON重复键并获取完整路径的解决方案
问题背景
需要解析包含重复键的嵌套JSON文件,输出重复键的完整层级路径及对应所有值。Python标准库的json.load/s会自动覆盖重复键,无法保留所有值;现有代码仅能检测重复键,但无法追踪其完整路径。
原生实现方案(无外部依赖)
通过自定义json.JSONDecoder,利用object_pairs_hook拦截字典解析过程,同时维护当前解析的层级路径,实现重复键的完整路径追踪。
完整代码
import json from typing import Dict, Any class DuplicateKeyDecoder(json.JSONDecoder): def __init__(self): super().__init__(object_pairs_hook=self.handle_object) self.current_path = [] self.duplicates = {} def handle_object(self, pairs): obj = {} for key, value in pairs: # 构建当前键的完整路径 full_path = " -> ".join(self.current_path + [key]) # 检测并记录重复键 if key in obj: if full_path not in self.duplicates: self.duplicates[full_path] = [obj[key], value] else: self.duplicates[full_path].append(value) else: obj[key] = value # 追踪嵌套字典的路径 if isinstance(value, dict): self.current_path.append(key) # 追踪数组中嵌套对象的路径(如 contacts[0].type) elif isinstance(value, list): for idx, item in enumerate(value): if isinstance(item, dict): array_node = f"{key}[{idx}]" self.current_path.append(array_node) self.decode(json.dumps(item)) self.current_path.pop() # 退出当前字典时,移除路径中的对应节点 if self.current_path and self.current_path[-1] in obj: self.current_path.pop() return obj def find_duplicate_keys(json_content: str) -> Dict[str, Any]: decoder = DuplicateKeyDecoder() decoder.decode(json_content) return decoder.duplicates # 测试示例 if __name__ == "__main__": sample_json = ''' { "name": "John", "age": 30, "address": { "street": "123 Main St", "city": "New York", "street": "321 Wall St" }, "contacts": [ { "type": "email", "value": "john@example.com" }, { "type": "phone", "value": "555-1234" }, { "type": "email", "value": "johndoe@example.com" } ], "age": 35 } ''' duplicates = find_duplicate_keys(sample_json) if duplicates: print("Duplicate keys found:") for path, values in duplicates.items(): # 格式化输出值(字符串加引号,数字直接输出) formatted_vals = ", ".join(f'"{v}"' if isinstance(v, str) else str(v) for v in values) print(f" {path} ({formatted_vals})") else: print("No duplicate keys found.")
代码说明
- 自定义解码器:
DuplicateKeyDecoder继承自标准库的json.JSONDecoder,通过object_pairs_hook拦截每个字典的键值对解析流程。 - 路径追踪:用
current_path列表维护当前解析的层级路径,进入嵌套字典时追加节点,退出时弹出节点。 - 重复键记录:解析每个键值对时,检查当前字典中是否已存在该键,若存在则记录完整路径及所有对应值。
- 数组支持:自动处理数组中嵌套对象的路径,格式如
contacts[0].type,但仅追踪同一对象内的重复键。
大文件适配方案(使用外部库)
若需处理超大JSON文件(无法一次性加载到内存),可使用ijson库实现流式解析,避免内存溢出。
示例代码
import ijson from typing import Dict, List def find_duplicates_large_json(file_path: str) -> Dict[str, List[Any]]: duplicates = {} current_path = [] current_obj = {} with open(file_path, 'r', encoding='utf-8') as f: for prefix, event, value in ijson.parse(f): if event == 'map_key': key = value full_path = " -> ".join(current_path + [key]) # 检查当前对象内的重复键 if key in current_obj: if full_path not in duplicates: duplicates[full_path] = [current_obj[key]] duplicates[full_path].append(value) current_obj[key] = value elif event == 'start_map': # 进入新字典,更新路径 if prefix: last_key = prefix.split('.')[-1] current_path.append(last_key) current_obj = {} elif event == 'end_map': # 退出字典,弹出路径节点 if current_path: current_path.pop() elif event == 'start_array': # 进入数组,记录数组节点 array_key = prefix.split('.')[-1] current_path.append(array_key) elif event == 'end_array': # 退出数组,弹出数组节点 if current_path: current_path.pop() elif event == 'item' and isinstance(value, dict): # 处理数组中的对象,更新路径为带索引的格式 idx = prefix.split('[')[-1].split(']')[0] current_path[-1] = f"{current_path[-1]}[{idx}]" return duplicates
使用说明
- 先安装
ijson:pip install ijson - 调用函数时传入JSON文件路径,即可流式解析并获取重复键信息。
内容的提问来源于stack exchange,提问作者Adam Cohen
相关产品推荐
相关产品推荐

