You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python的ijson流式处理大型JSON文件遇‘trailing garbage’错误求解决

解决ijson处理JSON Lines大文件的问题

问题背景

你尝试用ijson流式处理3GB的大型JSON文件,原代码如下:

with open('file.json', 'rb') as f:
    j = ijson.items(f, 'item')

for item in j:
    print('x')

运行时报trailing garbage错误,设置multiple_items=True后无报错但无输出。该文件为JSON Lines格式(每行一个独立JSON对象),结构示例:

{"_id":{"$oid":"6457879fd1187d621cbbba9c"},"sourceCC":"us",...}
{"_id":{"$oid":"6457879fd1187d621cbddd8a"},"sourceCC":"us",...}

原因分析

ijson默认解析的是标准单个顶级JSON结构(如单个对象或数组),而你的文件是多行独立JSON对象:

  1. 原代码中'item'路径期望JSON存在顶级数组,每个元素对应item,但你的文件无此结构,因此无匹配结果;
  2. 未设置multiple_items=True时,解析完第一个对象后,后续内容被判定为无效"垃圾数据",触发报错;
  3. 设置multiple_items=True后,ijson允许解析多个顶级对象,但'item'路径仍不匹配,所以无输出。

解决方案

方案一:逐行读取+json.loads(推荐,简单高效)

JSON Lines格式天然适合逐行处理,直接读取每行并用json.loads解析,内存占用低:

import json

with open('file.json', 'r', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        if not line:  # 跳过空行
            continue
        item = json.loads(line)
        print('x')
        # 在这里添加你的业务处理逻辑

方案二:用ijson的raw解析器处理连续对象

如果一定要用ijson,可以借助其底层解析器逐个识别对象的结束事件:

import ijson

with open('file.json', 'rb') as f:
    parser = ijson.parse(f)
    current_item = {}
    for prefix, event, value in parser:
        if event == 'start_map':
            current_item = {}
        elif event == 'map_key':
            current_key = value
        elif event in ('string', 'number', 'boolean', 'null'):
            current_item[current_key] = value
        elif event == 'end_map':
            print('x')
            # 处理current_item,比如写入数据库或做其他操作
            current_item = {}

内容的提问来源于stack exchange,提问作者Tim Vowden

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 01:04:52