使用ijson包读取大JSON文件时遭遇http.client.IncompleteRead错误
解决大JSON在线流式读取时的
IncompleteRead错误 首先咱们明确:这个错误不是文件过大导致的——毕竟你本地读取下载好的文件完全正常。问题出在HTTP流式传输的连接稳定性上,结合报错栈来看,是服务器端或网络中途断开了chunked编码的HTTP连接,导致数据读取不完整。
你的场景复盘
你用ijson处理Scryfall的1.5GB+全卡牌数据,在线流式处理的代码是:
response = requests.get("https://api.scryfall.com/bulk-data/all-cards") with urlopen(response.json()["download_uri"]) as all_cards: for card_object in ijson.items(all_cards, "item"): do_something_with(card_object)
但每次都会触发http.client.IncompleteRead错误,而本地读取下载后的文件却毫无问题:
with open("all-cards-20220408091307.json") as all_cards: for card_object in ijson.items(all_cards, "item"): do_something_with(card_object)
错误核心原因
这个报错的根源是HTTP连接在流式传输过程中被中断,常见触发因素包括:
- 服务器端超时:Scryfall的API可能对长时间保持的流式连接有超时限制,如果你的
do_something_with处理单条卡牌数据耗时较长,服务器会主动断开连接 - 网络不稳定:中间网络节点丢包、延迟过高,导致chunked传输的数据流断裂
urlopen的缺陷:标准库urlopen默认没有设置超时,反而可能因为连接闲置被服务器踢掉,或者遇到网络问题时一直挂着直到触发服务器端超时
针对性解决方案
方案1:改用requests流式下载+合理设置超时
requests库的流式处理比标准库urlopen更稳定,还能灵活设置连接和读取超时,避免连接被意外中断。修改后的代码如下:
import requests import ijson # 获取下载链接 response = requests.get("https://api.scryfall.com/bulk-data/all-cards") download_url = response.json()["download_uri"] # 开启流式下载,设置超时参数(连接超时10秒,读取超时30秒,可根据你的处理速度调整) with requests.get(download_url, stream=True, timeout=(10, 30)) as r: r.raise_for_status() # 先确保请求成功,避免下载无效数据 # 开启自动解压gzip(Scryfall的bulk数据默认是gzip压缩的) r.raw.decode_content = True # 直接用r.raw作为ijson的输入,流式处理 for card_object in ijson.items(r.raw, "item"): do_something_with(card_object)
方案2:增加重试机制应对网络波动
如果你的网络环境不太稳定,可以给请求加上重试逻辑,用tenacity库实现很方便:
import requests import ijson import http.client from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type # 定义重试装饰器:最多重试3次,间隔指数增长(2s、4s、8s),只重试网络相关异常 @retry( stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10), retry=retry_if_exception_type((requests.exceptions.RequestException, http.client.IncompleteRead)) ) def fetch_and_process_cards(): response = requests.get("https://api.scryfall.com/bulk-data/all-cards") download_url = response.json()["download_uri"] with requests.get(download_url, stream=True, timeout=(10, 30)) as r: r.raise_for_status() r.raw.decode_content = True for card_object in ijson.items(r.raw, "item"): do_something_with(card_object) # 执行任务 fetch_and_process_cards()
方案3:先完整下载再处理(最稳妥)
如果在线流式处理总是出问题,不如先把文件完整下载到本地,再用ijson处理——这是最不容易出网络相关问题的方式:
import requests import ijson # 第一步:获取下载链接并下载文件 response = requests.get("https://api.scryfall.com/bulk-data/all-cards") download_url = response.json()["download_uri"] local_file = "all-cards-latest.json" with requests.get(download_url, stream=True, timeout=(10, 30)) as r: r.raise_for_status() with open(local_file, 'wb') as f: # 分块下载,避免占用过多内存 for chunk in r.iter_content(chunk_size=8192): f.write(chunk) # 第二步:本地流式处理文件 with open(local_file, 'r') as all_cards: for card_object in ijson.items(all_cards, "item"): do_something_with(card_object)
内容的提问来源于stack exchange,提问作者Benjamin
相关产品推荐
相关产品推荐

