You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用ijson包读取大JSON文件时遭遇http.client.IncompleteRead错误

解决大JSON在线流式读取时的IncompleteRead错误

首先咱们明确:这个错误不是文件过大导致的——毕竟你本地读取下载好的文件完全正常。问题出在HTTP流式传输的连接稳定性上,结合报错栈来看,是服务器端或网络中途断开了chunked编码的HTTP连接,导致数据读取不完整。

你的场景复盘

你用ijson处理Scryfall的1.5GB+全卡牌数据,在线流式处理的代码是:

response = requests.get("https://api.scryfall.com/bulk-data/all-cards")
with urlopen(response.json()["download_uri"]) as all_cards:
    for card_object in ijson.items(all_cards, "item"):
        do_something_with(card_object)

但每次都会触发http.client.IncompleteRead错误,而本地读取下载后的文件却毫无问题:

with open("all-cards-20220408091307.json") as all_cards:
    for card_object in ijson.items(all_cards, "item"):
        do_something_with(card_object)

错误核心原因

这个报错的根源是HTTP连接在流式传输过程中被中断,常见触发因素包括:

  • 服务器端超时:Scryfall的API可能对长时间保持的流式连接有超时限制,如果你的do_something_with处理单条卡牌数据耗时较长,服务器会主动断开连接
  • 网络不稳定:中间网络节点丢包、延迟过高,导致chunked传输的数据流断裂
  • urlopen的缺陷:标准库urlopen默认没有设置超时,反而可能因为连接闲置被服务器踢掉,或者遇到网络问题时一直挂着直到触发服务器端超时

针对性解决方案

方案1:改用requests流式下载+合理设置超时

requests库的流式处理比标准库urlopen更稳定,还能灵活设置连接和读取超时,避免连接被意外中断。修改后的代码如下:

import requests
import ijson

# 获取下载链接
response = requests.get("https://api.scryfall.com/bulk-data/all-cards")
download_url = response.json()["download_uri"]

# 开启流式下载,设置超时参数(连接超时10秒,读取超时30秒,可根据你的处理速度调整)
with requests.get(download_url, stream=True, timeout=(10, 30)) as r:
    r.raise_for_status()  # 先确保请求成功,避免下载无效数据
    # 开启自动解压gzip(Scryfall的bulk数据默认是gzip压缩的)
    r.raw.decode_content = True
    # 直接用r.raw作为ijson的输入,流式处理
    for card_object in ijson.items(r.raw, "item"):
        do_something_with(card_object)

方案2:增加重试机制应对网络波动

如果你的网络环境不太稳定,可以给请求加上重试逻辑,用tenacity库实现很方便:

import requests
import ijson
import http.client
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

# 定义重试装饰器:最多重试3次,间隔指数增长(2s、4s、8s),只重试网络相关异常
@retry(
    stop=stop_after_attempt(3),
    wait=wait_exponential(multiplier=1, min=2, max=10),
    retry=retry_if_exception_type((requests.exceptions.RequestException, http.client.IncompleteRead))
)
def fetch_and_process_cards():
    response = requests.get("https://api.scryfall.com/bulk-data/all-cards")
    download_url = response.json()["download_uri"]
    
    with requests.get(download_url, stream=True, timeout=(10, 30)) as r:
        r.raise_for_status()
        r.raw.decode_content = True
        for card_object in ijson.items(r.raw, "item"):
            do_something_with(card_object)

# 执行任务
fetch_and_process_cards()

方案3:先完整下载再处理(最稳妥)

如果在线流式处理总是出问题,不如先把文件完整下载到本地,再用ijson处理——这是最不容易出网络相关问题的方式:

import requests
import ijson

# 第一步:获取下载链接并下载文件
response = requests.get("https://api.scryfall.com/bulk-data/all-cards")
download_url = response.json()["download_uri"]
local_file = "all-cards-latest.json"

with requests.get(download_url, stream=True, timeout=(10, 30)) as r:
    r.raise_for_status()
    with open(local_file, 'wb') as f:
        # 分块下载,避免占用过多内存
        for chunk in r.iter_content(chunk_size=8192):
            f.write(chunk)

# 第二步:本地流式处理文件
with open(local_file, 'r') as all_cards:
    for card_object in ijson.items(all_cards, "item"):
        do_something_with(card_object)

内容的提问来源于stack exchange,提问作者Benjamin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.28 18:18:13