You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PySpider爬取tesla.com遇HTTP 599错误,求解决及预防方法

解决PySpider爬取tesla.com时的HTTP 599错误及预防方案

问题描述

我用PySpider爬取tesla.com时持续出现报错:Exception: HTTP 599: HTTP/2 stream 0 was not closed cleanly: INTERNAL_ERROR (err 2),但用完全相同的代码爬取scrapy.org却能正常运行。我的代码如下:

from pyspider.libs.base_handler import *


class Handler(BaseHandler):
    crawl_config = {
    }

    @every(minutes=24 * 60)
    def on_start(self):
        self.crawl('https://www.tesla.com', callback=self.index_page, validate_cert=False)

    @config(age=10 * 24 * 60 * 60)
    def index_page(self, response):
        for each in response.doc('a[href^="http"]').items():
            self.crawl(each.attr.href, callback=self.detail_page, validate_cert=False)

    @config(priority=2)
    def detail_page(self, response):
        return {
            "url": response.url,
            "title": response.doc('title').text(),
        }

错误解决方法

  • 强制切换到HTTP/1.1协议
    Tesla服务器可能对HTTP/2协议存在兼容性问题,PySpider默认启用HTTP/2,在crawl_config中禁用HTTP/2并指定Connection头:

    crawl_config = {
        'headers': {
            'Connection': 'close'
        },
        'http2': False
    }
    
  • 模拟真实浏览器请求头
    服务器可能通过请求头识别出爬虫,添加常见的浏览器标识字段,让请求更接近真实用户:

    crawl_config = {
        'headers': {
            'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
            'Accept-Language': 'zh-CN,zh;q=0.9',
            'Connection': 'close'
        },
        'http2': False
    }
    
  • 增加请求延迟
    短时间内高频请求触发了服务器的反爬机制,在页面处理的装饰器中添加延迟参数:

    @config(age=10 * 24 * 60 * 60, delay=2)
    def index_page(self, response):
        for each in response.doc('a[href^="http"]').items():
            self.crawl(each.attr.href, callback=self.detail_page, validate_cert=False)
    

预防这类报错的通用方案

  • 优先禁用HTTP/2:对部分对HTTP/2支持不佳或有针对性反爬的网站,直接强制使用HTTP/1.1,避免协议层的兼容性问题。
  • 完善请求头模拟:除了User-Agent,还可以补充Accept、Referer、Cookie等字段,尽可能还原真实浏览器的请求特征。
  • 严格控制请求频率:通过delay设置请求间隔,或者搭配代理IP池分散请求来源,降低被服务器识别为爬虫的概率。
  • 启用自动重试机制:在crawl_config中配置重试次数和间隔,遇到临时网络或服务器错误时自动重试:
    crawl_config = {
        'retries': 3,
        'retry_interval': 5
    }
    
  • 分步测试爬取范围:不要一开始就大规模爬取全站,先测试单个首页,确认请求正常后再逐步扩展到子页面。

内容的提问来源于stack exchange,提问作者Kristin Chia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.23 05:42:34