如何为多台Scrapy爬虫实现proxy pool?已有动态更新代理数据库
Absolutely! This is exactly what Scrapy's downloader middlewares were built for—you can set up a global proxy middleware that pulls from your existing proxy database, handles rotation, and even invalidates dead proxies without touching a single line of your existing spiders. Here's a step-by-step solution tailored to your needs:
1. Build a Custom Downloader Middleware
Create a middleware class that integrates with your proxy database, manages proxy rotation, and handles failed requests by marking proxies as invalid. This middleware will intercept every request before it's sent, and every response/exception after.
from scrapy import signals from scrapy.downloadermiddlewares.httpproxy import HttpProxyMiddleware from scrapy.exceptions import ConnectTimeoutError, ConnectionRefusedError import your_db_module # Replace with your actual database handling code class GlobalProxyPoolMiddleware(HttpProxyMiddleware): def __init__(self, proxy_db): super().__init__() self.proxy_db = proxy_db self.proxy_list = self._fetch_valid_proxies() self.current_index = 0 # Refresh proxy list every 100 requests to keep it up-to-date self.refresh_threshold = 100 self.request_count = 0 @classmethod def from_crawler(cls, crawler): # Initialize your database connection here (pull config from settings if needed) proxy_db = your_db_module.ProxyDatabase() middleware = cls(proxy_db) # Clean up resources when spider closes crawler.signals.connect(middleware.spider_closed, signal=signals.spider_closed) return middleware def _fetch_valid_proxies(self): # Query your database for recently verified, active proxies return self.proxy_db.get_active_proxies() def _get_next_proxy(self): # Simple round-robin rotation; adjust to random/weighted if needed if not self.proxy_list: raise ValueError("No valid proxies available in the database") self.current_index = (self.current_index + 1) % len(self.proxy_list) return self.proxy_list[self.current_index] def process_request(self, request, spider): # Skip if the request already has a proxy (for spider-specific overrides) if 'proxy' in request.meta: return self.request_count += 1 # Refresh proxy list periodically to pick up new valid proxies if self.request_count % self.refresh_threshold == 0: self.proxy_list = self._fetch_valid_proxies() self.current_index = 0 proxy = self._get_next_proxy() # Set proxy URL (adjust for HTTPS if needed) request.meta['proxy'] = f"http://{proxy['ip_address']}:{proxy['port']}" # Add auth headers if your proxies require credentials if proxy.get('username') and proxy.get('password'): import base64 auth_str = base64.b64encode(f"{proxy['username']}:{proxy['password']}".encode()).decode() request.headers['Proxy-Authorization'] = f"Basic {auth_str}" def process_response(self, request, response, spider): # Check for proxy failure signs (adjust status codes/keywords to your use case) if response.status in [403, 502, 503] or 'captcha' in response.text.lower(): # Mark proxy as invalid in your database proxy_ip = request.meta['proxy'].split('//')[1].split(':')[0] self.proxy_db.mark_proxy_inactive(proxy_ip) # Rotate to a new proxy and retry the request new_proxy = self._get_next_proxy() request.meta['proxy'] = f"http://{new_proxy['ip_address']}:{new_proxy['port']}" return request.copy() return response def process_exception(self, request, exception, spider): # Handle connection-related errors that indicate a dead proxy if isinstance(exception, (ConnectTimeoutError, ConnectionRefusedError)): proxy_ip = request.meta['proxy'].split('//')[1].split(':')[0] self.proxy_db.mark_proxy_inactive(proxy_ip) new_proxy = self._get_next_proxy() request.meta['proxy'] = f"http://{new_proxy['ip_address']}:{new_proxy['port']}" return request.copy() def spider_closed(self, spider): # Optional: Clean up database connections here self.proxy_db.close_connection()
2. Enable the Middleware Globally
Open your project's settings.py and update the DOWNLOADER_MIDDLEWARES to enable your custom middleware and disable Scrapy's default proxy middleware to avoid conflicts:
DOWNLOADER_MIDDLEWARES = { # Your custom middleware (priority 700 runs before most default middlewares) 'your_project_name.middlewares.GlobalProxyPoolMiddleware': 700, # Disable Scrapy's default proxy middleware 'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware': None, }
3. Key Optimizations for Your Use Case
- Reduce Database Load: Instead of fetching proxies on every request, refresh the list periodically (like every 100 requests or every 10 minutes) to balance freshness and performance.
- Smart Proxy Selection: Replace round-robin with a strategy that prioritizes proxies with higher success rates or faster response times (store this data in your database!).
- Retry Configuration: Tweak Scrapy's built-in retry settings to align with your proxy strategy:
RETRY_TIMES = 3 # Allow 3 retries per request RETRY_HTTP_CODES = [403, 502, 503, 504, 429] # Retry on these status codes - Thread Safety: Ensure your database operations are thread-safe (use connection pools or thread-local connections) since Scrapy runs requests in parallel.
Once set up, all your existing spiders will automatically use the proxy pool—no code changes needed. The middleware will handle rotating proxies, marking dead ones in your database, and retrying failed requests seamlessly.
内容的提问来源于stack exchange,提问作者Paulo Cirino

