使用Python的ijson流式处理大型JSON文件遇‘trailing garbage’错误求解决
解决ijson处理JSON Lines大文件的问题
问题背景
你尝试用ijson流式处理3GB的大型JSON文件,原代码如下:
with open('file.json', 'rb') as f: j = ijson.items(f, 'item') for item in j: print('x')
运行时报trailing garbage错误,设置multiple_items=True后无报错但无输出。该文件为JSON Lines格式(每行一个独立JSON对象),结构示例:
{"_id":{"$oid":"6457879fd1187d621cbbba9c"},"sourceCC":"us",...} {"_id":{"$oid":"6457879fd1187d621cbddd8a"},"sourceCC":"us",...}
原因分析
ijson默认解析的是标准单个顶级JSON结构(如单个对象或数组),而你的文件是多行独立JSON对象:
- 原代码中
'item'路径期望JSON存在顶级数组,每个元素对应item,但你的文件无此结构,因此无匹配结果; - 未设置
multiple_items=True时,解析完第一个对象后,后续内容被判定为无效"垃圾数据",触发报错; - 设置
multiple_items=True后,ijson允许解析多个顶级对象,但'item'路径仍不匹配,所以无输出。
解决方案
方案一:逐行读取+json.loads(推荐,简单高效)
JSON Lines格式天然适合逐行处理,直接读取每行并用json.loads解析,内存占用低:
import json with open('file.json', 'r', encoding='utf-8') as f: for line in f: line = line.strip() if not line: # 跳过空行 continue item = json.loads(line) print('x') # 在这里添加你的业务处理逻辑
方案二:用ijson的raw解析器处理连续对象
如果一定要用ijson,可以借助其底层解析器逐个识别对象的结束事件:
import ijson with open('file.json', 'rb') as f: parser = ijson.parse(f) current_item = {} for prefix, event, value in parser: if event == 'start_map': current_item = {} elif event == 'map_key': current_key = value elif event in ('string', 'number', 'boolean', 'null'): current_item[current_key] = value elif event == 'end_map': print('x') # 处理current_item,比如写入数据库或做其他操作 current_item = {}
内容的提问来源于stack exchange,提问作者Tim Vowden
相关产品推荐
相关产品推荐

