You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用ijson高效解析大型gzip压缩JSON文件并避免‘trailing garbage’错误?

问题描述

处理一个1500万+行的gzip压缩JSON文件,格式为每行一个独立JSON对象,目标是高效提取review_text字段且不加载全量数据到内存。使用ijson解析时触发以下错误:

IncompleteJSONError: parse error: trailing garbage
votes": 16, "n_comments": 0} {"user_id": "8842281e1d1347389f
(right here) ------^

数据示例(每行一个独立JSON):

{"user_id": "8842281e1d1347389f2ab93d60773d4d", "book_id": "24375664", "review_id": "5cd416f3efc3f944fce4ce2db2290d5e", "rating": 5, "review_text": "Mind blowingly cool. Best science fiction I've read in some time...", "date_added": "Fri Aug 25 13:55:02 -0700 2017", "date_updated": "Mon Oct 09 08:55:59 -0700 2017", "read_at": "Sat Oct 07 00:00:00 -0700 2017", "started_at": "Sat Aug 26 00:00:00 -0700 2017", "n_votes": 16, "n_comments": 0}
{"user_id": "8842281e1d1347389f2ab93d60773d4d", "book_id": "18245960", "review_id": "dfdbb7b0eb5a7e4c26d59a937e2e5feb", "rating": 5, "review_text": "This is a special book. It started slow for about the first third...", "date_added": "Sun Jul 30 07:44:10 -0700 2017", "date_updated": "Wed Aug 30 00:00:26 -0700 2017", "read_at": "Sat Aug 26 12:05:52 -0700 2017", "started_at": "Tue Aug 15 13:23:18 -0700 2017", "n_votes": 28, "n_comments": 1}

已尝试的方法:

  • 直接用ijson解析:触发上述错误,原因是ijson默认解析JSON数组格式,但该文件是行分隔JSON(JSON Lines),多个对象直接拼接无数组包裹,解析第一个对象后剩余内容被判定为垃圾数据。
  • 逐行清理后解析JSON:效率不足,无法应对超大规模文件。

触发错误的代码:

import ijson
import pandas as pd
import gzip

review_texts = []

gzip_file_path = 'goodreads_dataset.json.gz'

with gzip.open(gzip_file_path, 'rt', encoding='utf-8') as f:
    objects = ijson.items(f, 'item')  # 错误假设文件是顶层数组格式

    for obj in objects:
        if 'review_text' in obj:
            review_texts.append(obj['review_text'])

df = pd.DataFrame(review_texts, columns=['review_text'])
df.to_pickle('reviews.pkl')

print(f"Saved {len(df)} review_text entries to 'reviews.pkl')

解决方案

针对JSON Lines格式的大文件,以下两种方式均能高效处理,且无需加载全量数据到内存:

方法1:ijson逐行解析(适配JSON Lines)

通过逐行读取并单独解析每行JSON对象,保留ijson的流式处理特性:

import ijson
import pandas as pd
import gzip
from io import StringIO

review_texts = []
gzip_file_path = 'goodreads_dataset.json.gz'
batch_size = 10000  # 批量写入阈值,控制内存占用

with gzip.open(gzip_file_path, 'rt', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        # 将单行字符串转为类文件对象供ijson解析
        line_io = StringIO(line)
        obj = next(ijson.items(line_io, ''))
        if 'review_text' in obj:
            review_texts.append(obj['review_text'])
        
        # 批量写入,避免内存溢出
        if len(review_texts) >= batch_size:
            pd.DataFrame(review_texts, columns=['review_text']).to_pickle('reviews_batch.pkl', mode='ab')
            review_texts = []

# 处理剩余的最后一批数据
if review_texts:
    pd.DataFrame(review_texts, columns=['review_text']).to_pickle('reviews_batch.pkl', mode='ab')

print("处理完成")

方法2:标准库json逐行解析(更轻量高效)

若无需ijson的路径查询功能,直接用标准库json逐行解析效率更高:

import json
import pandas as pd
import gzip

review_texts = []
gzip_file_path = 'goodreads_dataset.json.gz'
batch_size = 10000
output_file = 'reviews.pkl'

with gzip.open(gzip_file_path, 'rt', encoding='utf-8') as f:
    for line in f:
        line = line.strip()
        if not line:
            continue
        try:
            obj = json.loads(line)
            if 'review_text' in obj:
                review_texts.append(obj['review_text'])
        except json.JSONDecodeError as e:
            print(f"跳过损坏行: {e}")
            continue
        
        # 达到批量阈值时写入文件
        if len(review_texts) >= batch_size:
            df = pd.DataFrame(review_texts, columns=['review_text'])
            # 首次写入用wb模式,后续用ab追加
            df.to_pickle(output_file, mode='wb' if len(review_texts) == batch_size else 'ab')
            review_texts = []

# 写入最后一批剩余数据
if review_texts:
    pd.DataFrame(review_texts, columns=['review_text']).to_pickle(output_file, mode='ab')

print(f"保存完成,累计处理有效记录约{len(review_texts) + (batch_size * (len(review_texts)//batch_size))}条")

关键调整说明

  1. 格式适配:明确文件是JSON Lines而非标准JSON数组,必须逐行处理单个JSON对象,而非用ijson解析整个文件。
  2. 内存控制:通过批量写入机制,避免内存中堆积过多数据,适配1500万条记录的规模。
  3. 容错处理:添加JSON解析错误捕获,防止个别损坏行导致程序中断。

内容的提问来源于stack exchange,提问作者chebz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.18 14:39:50