You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ijson解析大JSON文件遇错误,如何跳过坏记录继续迭代?

解决ijson解析损坏JSON时跳过错误记录的问题

针对你遇到的问题,直接用ijson.items()会因为解析器抛出错误而终止整个迭代,要实现跳过损坏记录继续解析,需要借助ijson的低级事件驱动API手动处理解析流程,在遇到错误时定位到下一个有效元素的起始点,恢复解析。

具体实现代码

import ijson
import traceback

def safe_parse_items(file_obj, target_prefix):
    # 拆分目标前缀,比如'rows.item'拆分为父路径和子项路径
    parent_path, item_path = target_prefix.rsplit('.', 1)
    parser = ijson.parse(file_obj)
    current_path = []
    item_event_buffer = []
    is_processing_item = False

    while True:
        try:
            event, value = next(parser)
        except ijson.common.IncompleteJSONError:
            print("检测到损坏的JSON记录,尝试跳过...")
            # 读取文件直到找到数组元素分隔符','或数组结束符']'
            while True:
                char = file_obj.read(1)
                if not char:  # 文件读取完毕
                    return
                if char in (',', ']'):
                    # 将指针回退一个字符,留给新的解析器处理
                    file_obj.seek(file_obj.tell() - 1)
                    # 重置解析器和状态
                    parser = ijson.parse(file_obj)
                    current_path = []
                    is_processing_item = False
                    item_event_buffer = []
                    break
            continue
        except StopIteration:
            break  # 解析完成

        # 维护当前解析的路径
        if event.startswith('start_'):
            current_path.append(value)
        elif event.startswith('end_'):
            if current_path:
                current_path.pop()

        # 生成当前路径的字符串标识
        current_prefix = '.'.join(current_path)

        # 捕获目标项的所有事件,用于重建对象
        if current_prefix == target_prefix:
            if event == 'start_map':
                is_processing_item = True
                item_event_buffer = []
            elif event == 'end_map':
                is_processing_item = False
                # 从事件缓冲区重建完整的JSON对象
                item = next(ijson.items_from_events(iter(item_event_buffer)))
                yield item
            if is_processing_item:
                item_event_buffer.append((event, value))

# 使用示例
try:
    with open('your_large_file.json', 'r', encoding='utf-8') as json_file:
        # 替换为你的目标前缀'rows.item'
        for row in safe_parse_items(json_file, 'rows.item'):
            try:
                super_df = dependency_latest_version(row, version)
            except Exception as e:
                print(f"处理单条记录出错: {str(e)}")
except Exception as e:
    traceback.print_exc()

原理说明

  1. 低级API控制解析流程:放弃ijson.items()的高级封装,改用ijson.parse()获取原始解析事件(如start_map、end_map等),手动跟踪解析路径。
  2. 错误恢复逻辑:当遇到IncompleteJSONError时,逐字符读取文件直到找到数组元素的分隔符(,或]),回退文件指针后重置解析器,继续处理后续内容。
  3. 事件缓冲区重建对象:在解析到有效目标项时,缓存其所有事件,待事件结束后用ijson.items_from_events()重建完整的JSON对象,确保输出的每条记录都是有效的。

这种方法可以在流式处理大文件的同时,跳过损坏的记录,继续完成剩余数据的解析。

内容的提问来源于stack exchange,提问作者aditi19

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 01:05:31