如何在JSON重复键报错信息中添加行号?支持含注释空行的JSON
如何在JSON重复键报错中添加对应行号
问题说明
需求:当JSON包含存在重复键的字典时抛出错误。
需要解决的问题:如何在报错信息中添加该重复键所在的JSON行号?该JSON可能包含注释或空行,不想手动统计行数,寻求更优解决方案。
现有实现代码
重复键检查钩子
import json def dict_raise_on_duplicates(ordered_pairs): """Reject duplicate keys.""" d = {} for k, v in ordered_pairs: if k in d: raise ValueError("duplicate key: %r" % (k,)) else: d[k] = v return d
示例JSON内容
{ "fruit": "Apple", "size": "Large", "size": "Red" }
主函数
def main(): try: data = json.loads(file_content, object_pairs_hook=dict_raise_on_duplicates) except ValueError as e: print("Error: the JSON has syntax error: " + str(e)) exit(1)
解决方案
原生object_pairs_hook无法直接获取键的位置信息,我们可以通过自定义JSON解码器,结合字符索引到行号的映射来实现行号追踪。核心思路是:
- 先构建JSON字符串的字符索引与行号的映射表,覆盖所有换行符格式(
\n、\r、\r\n)。 - 重写JSON解码器的对象解析逻辑,记录每个键的起始字符索引,检查重复键时通过映射表找到对应行号。
完整实现代码
import json from json.decoder import JSONDecodeError, WHITESPACE def build_line_map(s): """构建字符索引到行号的映射表""" line_map = {} current_line = 1 for idx, char in enumerate(s): line_map[idx] = current_line if char == '\n': current_line += 1 elif char == '\r': # 处理单独的\r,跳过\r\n的情况 if idx + 1 >= len(s) or s[idx + 1] != '\n': current_line += 1 return line_map def dict_raise_on_duplicates(ordered_pairs, line_map): """检查重复键并抛出带行号的错误""" d = {} for k, v, key_start_idx in ordered_pairs: if k in d: line_num = line_map[key_start_idx] raise ValueError(f"duplicate key: {repr(k)} at line {line_num}") d[k] = v return d class LineAwareJSONDecoder(json.JSONDecoder): """支持追踪键行号的JSON解码器""" def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) # 重写对象解析方法 self.parse_object = self._parse_object_with_pos def _parse_object_with_pos(self, s, idx): """解析JSON对象并记录每个键的起始位置""" idx = WHITESPACE.match(s, idx).end() if s[idx] != '{': raise JSONDecodeError("Expecting '{'", s, idx) idx += 1 idx = WHITESPACE.match(s, idx).end() pairs = [] while True: idx = WHITESPACE.match(s, idx).end() # 处理对象结束符 if s[idx] == '}': idx += 1 return dict_raise_on_duplicates(pairs, self.line_map), idx # 处理键值对分隔符 if s[idx] == ',': idx += 1 idx = WHITESPACE.match(s, idx).end() # 记录键的起始索引 key_start_idx = idx # 解析键 key, idx = self.parse_string(s, key_start_idx) idx = WHITESPACE.match(s, idx).end() if s[idx] != ':': raise JSONDecodeError("Expecting ':'", s, idx) idx += 1 idx = WHITESPACE.match(s, idx).end() # 解析值 value, idx = self.raw_decode(s, idx) pairs.append((key, value, key_start_idx)) idx = WHITESPACE.match(s, idx).end() def main(file_content): # 构建行号映射 line_map = build_line_map(file_content) # 初始化自定义解码器 decoder = LineAwareJSONDecoder() decoder.line_map = line_map try: data = decoder.decode(file_content) print("JSON解析成功:", data) except ValueError as e: print(f"Error: the JSON has syntax error: {str(e)}") exit(1) # 测试用例 if __name__ == "__main__": test_json = ''' { "fruit": "Apple", "size": "Large", "size": "Red" } ''' main(test_json)
代码说明
build_line_map:遍历JSON字符串,为每个字符位置绑定对应的行号,兼容多种换行格式。LineAwareJSONDecoder:重写原生解码器的对象解析逻辑,在解析每个键时记录其起始字符索引,传递给重复键检查函数。dict_raise_on_duplicates:检查到重复键时,通过字符索引在映射表中找到行号,抛出包含行号的错误信息。
处理带注释的JSON
原生JSON规范不支持注释,如果你的JSON包含注释,需要先预处理移除注释,例如用正则表达式清理:
import re def remove_json_comments(s): """移除JSON中的//单行注释和/* */多行注释""" # 移除多行注释 s = re.sub(r'/\*.*?\*/', '', s, flags=re.DOTALL) # 移除单行注释 s = re.sub(r'//.*', '', s) return s # 使用时先清理注释 cleaned_content = remove_json_comments(file_content) main(cleaned_content)
内容的提问来源于stack exchange,提问作者Alex
相关产品推荐
相关产品推荐

