PySpider爬取tesla.com遇HTTP 599错误,求解决及预防方法
解决PySpider爬取tesla.com时的HTTP 599错误及预防方案
问题描述
我用PySpider爬取tesla.com时持续出现报错:Exception: HTTP 599: HTTP/2 stream 0 was not closed cleanly: INTERNAL_ERROR (err 2),但用完全相同的代码爬取scrapy.org却能正常运行。我的代码如下:
from pyspider.libs.base_handler import * class Handler(BaseHandler): crawl_config = { } @every(minutes=24 * 60) def on_start(self): self.crawl('https://www.tesla.com', callback=self.index_page, validate_cert=False) @config(age=10 * 24 * 60 * 60) def index_page(self, response): for each in response.doc('a[href^="http"]').items(): self.crawl(each.attr.href, callback=self.detail_page, validate_cert=False) @config(priority=2) def detail_page(self, response): return { "url": response.url, "title": response.doc('title').text(), }
错误解决方法
强制切换到HTTP/1.1协议
Tesla服务器可能对HTTP/2协议存在兼容性问题,PySpider默认启用HTTP/2,在crawl_config中禁用HTTP/2并指定Connection头:crawl_config = { 'headers': { 'Connection': 'close' }, 'http2': False }模拟真实浏览器请求头
服务器可能通过请求头识别出爬虫,添加常见的浏览器标识字段,让请求更接近真实用户:crawl_config = { 'headers': { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'zh-CN,zh;q=0.9', 'Connection': 'close' }, 'http2': False }增加请求延迟
短时间内高频请求触发了服务器的反爬机制,在页面处理的装饰器中添加延迟参数:@config(age=10 * 24 * 60 * 60, delay=2) def index_page(self, response): for each in response.doc('a[href^="http"]').items(): self.crawl(each.attr.href, callback=self.detail_page, validate_cert=False)
预防这类报错的通用方案
- 优先禁用HTTP/2:对部分对HTTP/2支持不佳或有针对性反爬的网站,直接强制使用HTTP/1.1,避免协议层的兼容性问题。
- 完善请求头模拟:除了User-Agent,还可以补充Accept、Referer、Cookie等字段,尽可能还原真实浏览器的请求特征。
- 严格控制请求频率:通过
delay设置请求间隔,或者搭配代理IP池分散请求来源,降低被服务器识别为爬虫的概率。 - 启用自动重试机制:在
crawl_config中配置重试次数和间隔,遇到临时网络或服务器错误时自动重试:crawl_config = { 'retries': 3, 'retry_interval': 5 } - 分步测试爬取范围:不要一开始就大规模爬取全站,先测试单个首页,确认请求正常后再逐步扩展到子页面。
内容的提问来源于stack exchange,提问作者Kristin Chia
相关产品推荐
相关产品推荐

