Scrapy如何等待JSON页面完全加载以避免解析失败?
解决Scrapy爬取JSON页面时的JSONDecodeError问题
问题说明
爬取JSON接口时偶尔出现解析失败错误,具体报错如下:
ERROR: Spider error processing <GET https://reqbin.com/echo/get/json/page/2> Traceback (most recent call last): File "/home/user/.local/lib/python3.8/site-packages/twisted/internet/defer.py", line 857, in _runCallbacks current.result = callback( # type: ignore[misc] File "/home/user/path/scraping.py", line 239, in parse_images jsonresponse = json.loads(response.text) File "/usr/lib/python3.8/json/__init__.py", line 357, in loads return _default_decoder.decode(s) File "/usr/lib/python3.8/json/decoder.py", line 337, in decode obj, end = self.raw_decode(s, idx=_w(s, 0).end()) File "/usr/lib/python3.8/json/decoder.py", line 353, in raw_decode obj, end = self.scan_once(s, idx) json.decoder.JSONDecodeError: Unterminated string starting at: line 1 column 48662 (char 48661)
已配置的Scrapy设置:
DOWNLOAD_DELAY = 5 DOWNLOAD_TIMEOUT = 600 DOWNLOAD_FAIL_ON_DATALOSS = False CONCURRENT_REQUESTS = 8
手动提取JSON内容可正常解析,怀疑是响应未完全接收导致解析失败。
解决方案
1. 改用Scrapy内置JSON解析方法
放弃json.loads(response.text),直接使用response.json()——Scrapy的该方法会自动处理编码校验、响应完整性检查,比手动调用json模块更适配爬虫场景:
def parse_images(self, response): try: jsonresponse = response.json() except ValueError as e: # 记录错误并触发重试 self.logger.error(f"JSON解析失败: {e}, URL: {response.url}") yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True)
2. 手动校验响应完整性
在解析函数中对比响应头的Content-Length和实际接收的字节数,不一致则重试:
def parse_images(self, response): content_length = response.headers.get('Content-Length') if content_length: expected_length = int(content_length.decode('utf-8')) actual_length = len(response.body) if actual_length != expected_length: self.logger.warning(f"响应不完整: 预期{expected_length}字节,实际{actual_length}字节,URL: {response.url}") yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True) return # 执行解析逻辑 try: jsonresponse = response.json() except ValueError as e: self.logger.error(f"JSON解析失败: {e}, URL: {response.url}") yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True)
3. 启用自动重试机制
在settings.py中配置重试参数,让Scrapy自动重试失败请求:
RETRY_ENABLED = True RETRY_TIMES = 3 # 重试次数 RETRY_HTTP_CODES = [500, 502, 503, 504, 408, 429]
4. 修正数据丢失检测设置
将DOWNLOAD_FAIL_ON_DATALOSS改为True,这样Scrapy检测到数据丢失时会直接标记请求失败,触发重试(之前设为False会忽略数据丢失,导致不完整响应进入解析环节):
DOWNLOAD_FAIL_ON_DATALOSS = True
5. 强制指定响应编码
部分场景下编码错误会导致response.text乱码,进而引发解析失败,可手动指定编码:
def parse_images(self, response): # 强制使用UTF-8编码解析 response.encoding = 'utf-8' try: jsonresponse = response.json() except ValueError as e: self.logger.error(f"JSON解析失败: {e}, URL: {response.url}") yield scrapy.Request(response.url, callback=self.parse_images, dont_filter=True)
内容的提问来源于stack exchange,提问作者Takamura
相关产品推荐
相关产品推荐

