You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python读取大型JSON数据集时Jupyter出现IOPub速率超限问题咨询

报错根因

你遇到的IOPub超限报错,核心原因是print(json.dumps(array, indent=4))这一步试图把完整的大容量数据集直接输出到Jupyter前端,超出了服务端默认的单请求输出数据量限制,数据读取环节本身是正常执行的。

正确处理方案
  • 优先选择抽样展示,不要全量输出
    验证数据结构、格式正确性时,仅输出统计信息和少量样本即可,完全不需要打印全量数据,示例代码如下:
    import json
    
    json_filename = 'test_data.json'
    array = {'foo': []}
    foo_list = array['foo']
    
    with open(json_filename, encoding="utf8") as file:
        for line in file:
            obj = json.loads(line)
            foo_list.append(obj)
    
    # 仅输出统计信息
    print(f"数据集总记录数:{len(foo_list)}")
    # 仅输出第一条数据验证格式
    print("单条数据样例:")
    print(json.dumps(foo_list[0], indent=4, ensure_ascii=False))
    
  • 重构后的结构化数据直接存本地文件,不要打印到前端
    如果你重编JSON结构是为了后续使用,直接写入本地文件即可,完全不需要输出到notebook页面:
    output_file = "structured_test_data.json"
    with open(output_file, "w", encoding="utf-8") as f:
        json.dump(array, f, indent=4, ensure_ascii=False)
    print(f"结构化数据已保存至:{output_file}")
    
  • 确需全量输出时临时调整配置(不推荐长期使用)
    启动Jupyter时添加参数调高速率限制即可:
    jupyter notebook --NotebookApp.iopub_data_rate_limit=1.0e10
    
  • 超大数据集优化处理建议
    针对Twitter这类行式存储的JSON数据集,可以直接用pandas读取生成结构化表格,无需手动解析拼接,内存占用更低、处理效率更高:
    import pandas as pd
    # 直接读取每行一个JSON对象的原始数据集
    df = pd.read_json("test_data.json", lines=True)
    # 查看前5行数据验证格式
    print(df.head())
    
    如果数据集体积远大于本地内存,不需要全量加载的场景下优先选择逐行处理逻辑,读一行处理一行,不要把所有数据都存入内存列表。

内容的提问来源于stack exchange,提问作者Ravi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 15:27:03