如何在Scrapy中实现HTTP 200不完整页面的请求重试?
Scrapy页面加载不完整问题的解决方法
一、先修复RetryMiddleware不生效的问题
默认RetryMiddleware的优先级是500,你之前设成800可能被其他中间件拦截,先把优先级改回默认值,同时检查配置是否正确:
# settings.py RETRY_ENABLED = True RETRY_TIMES = 3 # 最多重试3次 RETRY_HTTP_CODES = [500, 502, 503, 504, 408, 429] DOWNLOADER_MIDDLEWARES = { 'scrapy.downloadermiddlewares.retry.RetryMiddleware': 500, # 其他中间件保持原有配置 }
注意:默认重试只针对HTTP错误(比如5xx、408),你遇到的是页面加载成功但数据缺失,属于内容错误,默认逻辑不会触发重试,得自定义重试条件。
二、自定义重试逻辑(针对内容缺失)
方法1:写个自定义中间件
继承官方RetryMiddleware,重写process_response方法,判断页面是否有你需要的数据,没有就触发重试:
# middlewares.py from scrapy.downloadermiddlewares.retry import RetryMiddleware class CustomContentRetryMiddleware(RetryMiddleware): def process_response(self, request, response, spider): # 只针对产品页做检查(根据你的URL规则调整判断条件) if 'product' in request.url: # 检查关键字段是否存在 script_price = response.xpath('//script[contains(text(), "script_price")]/text()').get() if not script_price: reason = "页面内容不完整,缺少script_price字段" return self._retry(request, reason, spider) or response # 其他情况走默认重试逻辑 return super().process_response(request, response, spider)
然后在settings里替换默认的RetryMiddleware:
# settings.py DOWNLOADER_MIDDLEWARES = { '你的项目名.middlewares.CustomContentRetryMiddleware': 500, 'scrapy.downloadermiddlewares.retry.RetryMiddleware': None, # 禁用默认的 }
方法2:在爬虫里手动触发重试
不想写中间件的话,直接在产品页的解析函数里判断,数据缺失就抛出RetryRequest:
# spiders/your_spider.py from scrapy.downloadermiddlewares.retry import RetryRequest def parse_product(self, response): script_price = response.xpath('//script[contains(text(), "script_price")]/text()').get() if not script_price: # 控制重试次数,避免无限循环 retry_count = response.meta.get('retry_times', 0) if retry_count < self.settings.get('RETRY_TIMES'): yield RetryRequest( response.url, meta={'retry_times': retry_count + 1}, dont_filter=True, # 避免被去重规则过滤 callback=self.parse_product ) return self.logger.error(f"页面重试3次仍失败: {response.url}") return # 正常提取数据的逻辑 try: price = script_price.split('"')[1] # 其他字段提取... yield { 'price': price, # 其他字段 } except AttributeError as e: self.logger.error(f"提取数据出错: {response.url}, 错误: {str(e)}")
三、动态页面用等待加载解决
如果页面是JS动态渲染的,单纯重试没用,得等元素加载完成,推荐用Scrapy-Playwright:
- 先安装:
pip install scrapy-playwright - 配置settings:
# settings.py DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = {"headless": True}
- 在爬虫里给分类页和产品页设不同的等待条件:
# spiders/your_spider.py import scrapy def start_requests(self): # 分类页等待分类列表加载完成(替换成你页面的实际元素选择器) yield scrapy.Request( url='你的分类页URL', callback=self.parse_category, meta={ "playwright": True, "playwright_wait_for": "div.category-item-list", # 等待分类元素出现 "playwright_include_page": True, } ) async def parse_category(self, response): page = response.meta["playwright_page"] await page.close() # 提取产品链接 product_links = response.xpath('//a[@class="product-link"]/@href').getall() for link in product_links: # 产品页等待包含script_price的脚本加载完成 yield scrapy.Request( url=response.urljoin(link), callback=self.parse_product, meta={ "playwright": True, "playwright_wait_for": 'script:contains("script_price")', "playwright_include_page": True, } ) async def parse_product(self, response): page = response.meta["playwright_page"] await page.close() # 这里提取数据时,页面已经加载完成 script_price = response.xpath('//script[contains(text(), "script_price")]/text()').get() if script_price: price = script_price.split('"')[1] yield {'price': price}
四、其他优化小技巧
- 加下载延迟:
DOWNLOAD_DELAY = 2,避免请求太频繁被网站限流,导致页面加载不全。 - 禁用缓存:
HTTPCACHE_ENABLED = False,确保每次请求都是最新的页面。 - 开DEBUG日志:
LOG_LEVEL = 'DEBUG',查看重试是否触发,以及页面响应内容,方便排查问题。
内容的提问来源于stack exchange,提问作者Square
相关产品推荐
相关产品推荐

