You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在JSON重复键报错信息中添加行号?支持含注释空行的JSON

如何在JSON重复键报错中添加对应行号

问题说明

需求:当JSON包含存在重复键的字典时抛出错误。
需要解决的问题:如何在报错信息中添加该重复键所在的JSON行号?该JSON可能包含注释或空行,不想手动统计行数,寻求更优解决方案。

现有实现代码

重复键检查钩子

import json
def dict_raise_on_duplicates(ordered_pairs):
    """Reject duplicate keys."""
    d = {}
    for k, v in ordered_pairs:
        if k in d:
           raise ValueError("duplicate key: %r" % (k,))
        else:
           d[k] = v
    return d

示例JSON内容

{
        "fruit": "Apple",
        "size": "Large",
        "size": "Red"
       }

主函数

def main():
    try:
        data = json.loads(file_content, object_pairs_hook=dict_raise_on_duplicates)
    except ValueError as e:
        print("Error: the JSON has syntax error: " + str(e))
        exit(1)

解决方案

原生object_pairs_hook无法直接获取键的位置信息,我们可以通过自定义JSON解码器,结合字符索引到行号的映射来实现行号追踪。核心思路是:

  1. 先构建JSON字符串的字符索引与行号的映射表,覆盖所有换行符格式(\n、\r、\r\n)。
  2. 重写JSON解码器的对象解析逻辑,记录每个键的起始字符索引,检查重复键时通过映射表找到对应行号。

完整实现代码

import json
from json.decoder import JSONDecodeError, WHITESPACE

def build_line_map(s):
    """构建字符索引到行号的映射表"""
    line_map = {}
    current_line = 1
    for idx, char in enumerate(s):
        line_map[idx] = current_line
        if char == '\n':
            current_line += 1
        elif char == '\r':
            # 处理单独的\r,跳过\r\n的情况
            if idx + 1 >= len(s) or s[idx + 1] != '\n':
                current_line += 1
    return line_map

def dict_raise_on_duplicates(ordered_pairs, line_map):
    """检查重复键并抛出带行号的错误"""
    d = {}
    for k, v, key_start_idx in ordered_pairs:
        if k in d:
            line_num = line_map[key_start_idx]
            raise ValueError(f"duplicate key: {repr(k)} at line {line_num}")
        d[k] = v
    return d

class LineAwareJSONDecoder(json.JSONDecoder):
    """支持追踪键行号的JSON解码器"""
    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        # 重写对象解析方法
        self.parse_object = self._parse_object_with_pos

    def _parse_object_with_pos(self, s, idx):
        """解析JSON对象并记录每个键的起始位置"""
        idx = WHITESPACE.match(s, idx).end()
        if s[idx] != '{':
            raise JSONDecodeError("Expecting '{'", s, idx)
        idx += 1
        idx = WHITESPACE.match(s, idx).end()
        
        pairs = []
        while True:
            idx = WHITESPACE.match(s, idx).end()
            # 处理对象结束符
            if s[idx] == '}':
                idx += 1
                return dict_raise_on_duplicates(pairs, self.line_map), idx
            # 处理键值对分隔符
            if s[idx] == ',':
                idx += 1
                idx = WHITESPACE.match(s, idx).end()
            
            # 记录键的起始索引
            key_start_idx = idx
            # 解析键
            key, idx = self.parse_string(s, key_start_idx)
            idx = WHITESPACE.match(s, idx).end()
            
            if s[idx] != ':':
                raise JSONDecodeError("Expecting ':'", s, idx)
            idx += 1
            idx = WHITESPACE.match(s, idx).end()
            
            # 解析值
            value, idx = self.raw_decode(s, idx)
            pairs.append((key, value, key_start_idx))
            
            idx = WHITESPACE.match(s, idx).end()

def main(file_content):
    # 构建行号映射
    line_map = build_line_map(file_content)
    # 初始化自定义解码器
    decoder = LineAwareJSONDecoder()
    decoder.line_map = line_map
    
    try:
        data = decoder.decode(file_content)
        print("JSON解析成功:", data)
    except ValueError as e:
        print(f"Error: the JSON has syntax error: {str(e)}")
        exit(1)

# 测试用例
if __name__ == "__main__":
    test_json = '''
   {
        "fruit": "Apple",
        "size": "Large",
        "size": "Red"
       }
'''
    main(test_json)

代码说明

  • build_line_map:遍历JSON字符串,为每个字符位置绑定对应的行号,兼容多种换行格式。
  • LineAwareJSONDecoder:重写原生解码器的对象解析逻辑,在解析每个键时记录其起始字符索引,传递给重复键检查函数。
  • dict_raise_on_duplicates:检查到重复键时,通过字符索引在映射表中找到行号,抛出包含行号的错误信息。

处理带注释的JSON

原生JSON规范不支持注释,如果你的JSON包含注释,需要先预处理移除注释,例如用正则表达式清理:

import re

def remove_json_comments(s):
    """移除JSON中的//单行注释和/* */多行注释"""
    # 移除多行注释
    s = re.sub(r'/\*.*?\*/', '', s, flags=re.DOTALL)
    # 移除单行注释
    s = re.sub(r'//.*', '', s)
    return s

# 使用时先清理注释
cleaned_content = remove_json_comments(file_content)
main(cleaned_content)

内容的提问来源于stack exchange,提问作者Alex

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 02:45:18