You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

为JustEat数据爬虫添加网络故障安全机制的技术需求

Got it, let's break down how to add the fault tolerance you need for your JustEat spider. You want two key things: keep retrying failed URLs until your network comes back and track which URLs couldn't be fetched due to network issues. Here's a step-by-step implementation:


1. Configure Scrapy's Core Retry Settings for Persistent Retries

Scrapy has built-in retry middleware, but we need to tweak it to prioritize failed requests and retry them almost indefinitely (until your network recovers). Add these settings to your settings.py:

# Enable retry middleware
RETRY_ENABLED = True

# Set an extremely high retry count (effectively infinite for practical use)
RETRY_TIMES = 99999

# Retry on common network/server errors
RETRY_HTTP_CODES = [500, 502, 503, 504, 408, 429]

# Make failed requests jump to the front of the queue
SCHEDULER_PRIORITY_QUEUE = 'scrapy.pqueues.DownloaderAwarePriorityQueue'
RETRY_PRIORITY_ADJUST = -1  # Lower priority number = higher execution priority

This setup ensures that any request failing due to network issues gets retried immediately before new requests are processed, and keeps trying until your connection is restored.


2. Custom Middleware to Log Permanently Failed URLs

We'll build a middleware that tracks URLs that couldn't be fetched even after all retries, and saves them to a text file.

First, create a middlewares.py file in your Scrapy project (if it doesn't exist) and add this code:

from scrapy import signals
import os

class FailedUrlLoggerMiddleware:
    def __init__(self):
        self.failed_urls_path = 'failed_justeat_urls.txt'
        # Optional: Clear the log file when the spider starts
        if os.path.exists(self.failed_urls_path):
            os.remove(self.failed_urls_path)

    @classmethod
    def from_crawler(cls, crawler):
        middleware = cls()
        crawler.signals.connect(middleware.on_request_failed, signal=signals.request_failed)
        crawler.signals.connect(middleware.on_spider_close, signal=signals.spider_closed)
        return middleware

    def on_request_failed(self, failure, request, spider):
        # Only log URLs that have exhausted all retry attempts
        retry_count = request.meta.get('retry_times', 0)
        if retry_count >= spider.settings.get('RETRY_TIMES'):
            error_details = f"URL: {request.url} | Error: {str(failure.value)}\n"
            with open(self.failed_urls_path, 'a') as log_file:
                log_file.write(error_details)
            spider.logger.error(f"Permanent failure for URL: {request.url}")

    def on_spider_close(self, spider):
        spider.logger.info(f"Failed URLs saved to: {self.failed_urls_path}")

Then enable this middleware in settings.py (replace your_project_name with your actual project name):

DOWNLOADER_MIDDLEWARES = {
    'your_project_name.middlewares.FailedUrlLoggerMiddleware': 543,
}

Your parse_takeaway_links method uses Selenium, which can throw network errors that Scrapy's default retry won't catch. We'll add exception handling to retry these requests too:

Modify the parse_takeaway_links method in your spider:

from selenium.common.exceptions import WebDriverException, TimeoutException
import time

def parse_takeaway_links(self, response):
    driver = None
    try:
        driver = webdriver.Chrome()
        driver.get(response.url)
        
        # Scroll to load all content (add a small sleep to avoid missing content)
        for _ in range(1, 100):
            driver.find_element_by_tag_name('html').send_keys(Keys.END)
            time.sleep(0.5)
        
        soup = BeautifulSoup(driver.page_source, "html.parser")
        takeaway_links = soup.find_all('a', class_='c-listing-item-link u-clearfix', href=True)
        
        for link in takeaway_links:
            full_url = 'https://www.just-eat.co.uk' + link['href']
            yield scrapy.Request(full_url, callback=self.parse_info)
    
    except (WebDriverException, TimeoutException) as e:
        self.logger.error(f"Selenium failed to load {response.url}: {str(e)}")
        # Retry the current page with high priority
        yield scrapy.Request(
            response.url,
            callback=self.parse_takeaway_links,
            dont_filter=True,  # Bypass duplicate filtering for retries
            priority=10
        )
    
    finally:
        if driver:
            driver.quit()  # Properly clean up the browser instance

This catches Selenium-specific network/timeouts and retries the page, ensuring you don't skip locations just because Chrome couldn't load them temporarily.


Final Notes

  • Infinite Retries: The high RETRY_TIMES value means the spider will keep retrying failed URLs until you manually stop it or your network recovers. If you want a hard limit later, just lower this number.
  • Failed URLs Log: The failed_justeat_urls.txt file will only contain URLs that couldn't be fetched after all retries, so you can easily revisit them later.
  • Resource Cleanup: Using driver.quit() instead of driver.close() ensures Selenium doesn't leave orphaned browser processes running.

内容的提问来源于stack exchange,提问作者Aman Praveen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 18:37:30