You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy如何等待JSON页面完全加载以避免解析失败?

解决Scrapy爬取JSON页面时的JSONDecodeError问题

问题说明

爬取JSON接口时偶尔出现解析失败错误,具体报错如下:

ERROR: Spider error processing <GET https://reqbin.com/echo/get/json/page/2>
Traceback (most recent call last):
  File "/home/user/.local/lib/python3.8/site-packages/twisted/internet/defer.py", line 857, in _runCallbacks
    current.result = callback( # type: ignore[misc]
  File "/home/user/path/scraping.py", line 239, in parse_images
    jsonresponse = json.loads(response.text)
  File "/usr/lib/python3.8/json/__init__.py", line 357, in loads
    return _default_decoder.decode(s)
  File "/usr/lib/python3.8/json/decoder.py", line 337, in decode
    obj, end = self.raw_decode(s, idx=_w(s, 0).end())
  File "/usr/lib/python3.8/json/decoder.py", line 353, in raw_decode
    obj, end = self.scan_once(s, idx)
json.decoder.JSONDecodeError: Unterminated string starting at: line 1 column 48662 (char 48661)

已配置的Scrapy设置:

DOWNLOAD_DELAY = 5
DOWNLOAD_TIMEOUT = 600
DOWNLOAD_FAIL_ON_DATALOSS = False
CONCURRENT_REQUESTS = 8

手动提取JSON内容可正常解析,怀疑是响应未完全接收导致解析失败。

解决方案

1. 改用Scrapy内置JSON解析方法

放弃json.loads(response.text),直接使用response.json()——Scrapy的该方法会自动处理编码校验、响应完整性检查,比手动调用json模块更适配爬虫场景:

def parse_images(self, response):
    try:
        jsonresponse = response.json()
    except ValueError as e:
        # 记录错误并触发重试
        self.logger.error(f"JSON解析失败: {e}, URL: {response.url}")
        yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True)

2. 手动校验响应完整性

在解析函数中对比响应头的Content-Length和实际接收的字节数,不一致则重试:

def parse_images(self, response):
    content_length = response.headers.get('Content-Length')
    if content_length:
        expected_length = int(content_length.decode('utf-8'))
        actual_length = len(response.body)
        if actual_length != expected_length:
            self.logger.warning(f"响应不完整: 预期{expected_length}字节,实际{actual_length}字节,URL: {response.url}")
            yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True)
            return
    # 执行解析逻辑
    try:
        jsonresponse = response.json()
    except ValueError as e:
        self.logger.error(f"JSON解析失败: {e}, URL: {response.url}")
        yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True)

3. 启用自动重试机制

在settings.py中配置重试参数,让Scrapy自动重试失败请求:

RETRY_ENABLED = True
RETRY_TIMES = 3  # 重试次数
RETRY_HTTP_CODES = [500, 502, 503, 504, 408, 429]

4. 修正数据丢失检测设置

将DOWNLOAD_FAIL_ON_DATALOSS改为True,这样Scrapy检测到数据丢失时会直接标记请求失败,触发重试(之前设为False会忽略数据丢失,导致不完整响应进入解析环节):

DOWNLOAD_FAIL_ON_DATALOSS = True

5. 强制指定响应编码

部分场景下编码错误会导致response.text乱码,进而引发解析失败,可手动指定编码:

def parse_images(self, response):
    # 强制使用UTF-8编码解析
    response.encoding = 'utf-8'
    try:
        jsonresponse = response.json()
    except ValueError as e:
        self.logger.error(f"JSON解析失败: {e}, URL: {response.url}")
        yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True)

内容的提问来源于stack exchange,提问作者Takamura

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.21 18:42:40