为JustEat数据爬虫添加网络故障安全机制的技术需求
Got it, let's break down how to add the fault tolerance you need for your JustEat spider. You want two key things: keep retrying failed URLs until your network comes back and track which URLs couldn't be fetched due to network issues. Here's a step-by-step implementation:
1. Configure Scrapy's Core Retry Settings for Persistent Retries
Scrapy has built-in retry middleware, but we need to tweak it to prioritize failed requests and retry them almost indefinitely (until your network recovers). Add these settings to your settings.py:
# Enable retry middleware RETRY_ENABLED = True # Set an extremely high retry count (effectively infinite for practical use) RETRY_TIMES = 99999 # Retry on common network/server errors RETRY_HTTP_CODES = [500, 502, 503, 504, 408, 429] # Make failed requests jump to the front of the queue SCHEDULER_PRIORITY_QUEUE = 'scrapy.pqueues.DownloaderAwarePriorityQueue' RETRY_PRIORITY_ADJUST = -1 # Lower priority number = higher execution priority
This setup ensures that any request failing due to network issues gets retried immediately before new requests are processed, and keeps trying until your connection is restored.
2. Custom Middleware to Log Permanently Failed URLs
We'll build a middleware that tracks URLs that couldn't be fetched even after all retries, and saves them to a text file.
First, create a middlewares.py file in your Scrapy project (if it doesn't exist) and add this code:
from scrapy import signals import os class FailedUrlLoggerMiddleware: def __init__(self): self.failed_urls_path = 'failed_justeat_urls.txt' # Optional: Clear the log file when the spider starts if os.path.exists(self.failed_urls_path): os.remove(self.failed_urls_path) @classmethod def from_crawler(cls, crawler): middleware = cls() crawler.signals.connect(middleware.on_request_failed, signal=signals.request_failed) crawler.signals.connect(middleware.on_spider_close, signal=signals.spider_closed) return middleware def on_request_failed(self, failure, request, spider): # Only log URLs that have exhausted all retry attempts retry_count = request.meta.get('retry_times', 0) if retry_count >= spider.settings.get('RETRY_TIMES'): error_details = f"URL: {request.url} | Error: {str(failure.value)}\n" with open(self.failed_urls_path, 'a') as log_file: log_file.write(error_details) spider.logger.error(f"Permanent failure for URL: {request.url}") def on_spider_close(self, spider): spider.logger.info(f"Failed URLs saved to: {self.failed_urls_path}")
Then enable this middleware in settings.py (replace your_project_name with your actual project name):
DOWNLOADER_MIDDLEWARES = { 'your_project_name.middlewares.FailedUrlLoggerMiddleware': 543, }
3. Fix Selenium-Related Network Failures
Your parse_takeaway_links method uses Selenium, which can throw network errors that Scrapy's default retry won't catch. We'll add exception handling to retry these requests too:
Modify the parse_takeaway_links method in your spider:
from selenium.common.exceptions import WebDriverException, TimeoutException import time def parse_takeaway_links(self, response): driver = None try: driver = webdriver.Chrome() driver.get(response.url) # Scroll to load all content (add a small sleep to avoid missing content) for _ in range(1, 100): driver.find_element_by_tag_name('html').send_keys(Keys.END) time.sleep(0.5) soup = BeautifulSoup(driver.page_source, "html.parser") takeaway_links = soup.find_all('a', class_='c-listing-item-link u-clearfix', href=True) for link in takeaway_links: full_url = 'https://www.just-eat.co.uk' + link['href'] yield scrapy.Request(full_url, callback=self.parse_info) except (WebDriverException, TimeoutException) as e: self.logger.error(f"Selenium failed to load {response.url}: {str(e)}") # Retry the current page with high priority yield scrapy.Request( response.url, callback=self.parse_takeaway_links, dont_filter=True, # Bypass duplicate filtering for retries priority=10 ) finally: if driver: driver.quit() # Properly clean up the browser instance
This catches Selenium-specific network/timeouts and retries the page, ensuring you don't skip locations just because Chrome couldn't load them temporarily.
Final Notes
- Infinite Retries: The high
RETRY_TIMESvalue means the spider will keep retrying failed URLs until you manually stop it or your network recovers. If you want a hard limit later, just lower this number. - Failed URLs Log: The
failed_justeat_urls.txtfile will only contain URLs that couldn't be fetched after all retries, so you can easily revisit them later. - Resource Cleanup: Using
driver.quit()instead ofdriver.close()ensures Selenium doesn't leave orphaned browser processes running.
内容的提问来源于stack exchange,提问作者Aman Praveen

