You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为多台Scrapy爬虫实现proxy pool?已有动态更新代理数据库

Absolutely! This is exactly what Scrapy's downloader middlewares were built for—you can set up a global proxy middleware that pulls from your existing proxy database, handles rotation, and even invalidates dead proxies without touching a single line of your existing spiders. Here's a step-by-step solution tailored to your needs:

Global Proxy Pool Middleware for Scrapy

1. Build a Custom Downloader Middleware

Create a middleware class that integrates with your proxy database, manages proxy rotation, and handles failed requests by marking proxies as invalid. This middleware will intercept every request before it's sent, and every response/exception after.

from scrapy import signals
from scrapy.downloadermiddlewares.httpproxy import HttpProxyMiddleware
from scrapy.exceptions import ConnectTimeoutError, ConnectionRefusedError
import your_db_module  # Replace with your actual database handling code

class GlobalProxyPoolMiddleware(HttpProxyMiddleware):
    def __init__(self, proxy_db):
        super().__init__()
        self.proxy_db = proxy_db
        self.proxy_list = self._fetch_valid_proxies()
        self.current_index = 0
        # Refresh proxy list every 100 requests to keep it up-to-date
        self.refresh_threshold = 100
        self.request_count = 0

    @classmethod
    def from_crawler(cls, crawler):
        # Initialize your database connection here (pull config from settings if needed)
        proxy_db = your_db_module.ProxyDatabase()
        middleware = cls(proxy_db)
        # Clean up resources when spider closes
        crawler.signals.connect(middleware.spider_closed, signal=signals.spider_closed)
        return middleware

    def _fetch_valid_proxies(self):
        # Query your database for recently verified, active proxies
        return self.proxy_db.get_active_proxies()

    def _get_next_proxy(self):
        # Simple round-robin rotation; adjust to random/weighted if needed
        if not self.proxy_list:
            raise ValueError("No valid proxies available in the database")
        self.current_index = (self.current_index + 1) % len(self.proxy_list)
        return self.proxy_list[self.current_index]

    def process_request(self, request, spider):
        # Skip if the request already has a proxy (for spider-specific overrides)
        if 'proxy' in request.meta:
            return

        self.request_count += 1
        # Refresh proxy list periodically to pick up new valid proxies
        if self.request_count % self.refresh_threshold == 0:
            self.proxy_list = self._fetch_valid_proxies()
            self.current_index = 0

        proxy = self._get_next_proxy()
        # Set proxy URL (adjust for HTTPS if needed)
        request.meta['proxy'] = f"http://{proxy['ip_address']}:{proxy['port']}"
        # Add auth headers if your proxies require credentials
        if proxy.get('username') and proxy.get('password'):
            import base64
            auth_str = base64.b64encode(f"{proxy['username']}:{proxy['password']}".encode()).decode()
            request.headers['Proxy-Authorization'] = f"Basic {auth_str}"

    def process_response(self, request, response, spider):
        # Check for proxy failure signs (adjust status codes/keywords to your use case)
        if response.status in [403, 502, 503] or 'captcha' in response.text.lower():
            # Mark proxy as invalid in your database
            proxy_ip = request.meta['proxy'].split('//')[1].split(':')[0]
            self.proxy_db.mark_proxy_inactive(proxy_ip)
            # Rotate to a new proxy and retry the request
            new_proxy = self._get_next_proxy()
            request.meta['proxy'] = f"http://{new_proxy['ip_address']}:{new_proxy['port']}"
            return request.copy()
        return response

    def process_exception(self, request, exception, spider):
        # Handle connection-related errors that indicate a dead proxy
        if isinstance(exception, (ConnectTimeoutError, ConnectionRefusedError)):
            proxy_ip = request.meta['proxy'].split('//')[1].split(':')[0]
            self.proxy_db.mark_proxy_inactive(proxy_ip)
            new_proxy = self._get_next_proxy()
            request.meta['proxy'] = f"http://{new_proxy['ip_address']}:{new_proxy['port']}"
            return request.copy()

    def spider_closed(self, spider):
        # Optional: Clean up database connections here
        self.proxy_db.close_connection()

2. Enable the Middleware Globally

Open your project's settings.py and update the DOWNLOADER_MIDDLEWARES to enable your custom middleware and disable Scrapy's default proxy middleware to avoid conflicts:

DOWNLOADER_MIDDLEWARES = {
    # Your custom middleware (priority 700 runs before most default middlewares)
    'your_project_name.middlewares.GlobalProxyPoolMiddleware': 700,
    # Disable Scrapy's default proxy middleware
    'scrapy.downloadermiddlewares.httpproxy.HttpProxyMiddleware': None,
}

3. Key Optimizations for Your Use Case

  • Reduce Database Load: Instead of fetching proxies on every request, refresh the list periodically (like every 100 requests or every 10 minutes) to balance freshness and performance.
  • Smart Proxy Selection: Replace round-robin with a strategy that prioritizes proxies with higher success rates or faster response times (store this data in your database!).
  • Retry Configuration: Tweak Scrapy's built-in retry settings to align with your proxy strategy:
    RETRY_TIMES = 3  # Allow 3 retries per request
    RETRY_HTTP_CODES = [403, 502, 503, 504, 429]  # Retry on these status codes
    
  • Thread Safety: Ensure your database operations are thread-safe (use connection pools or thread-local connections) since Scrapy runs requests in parallel.

Once set up, all your existing spiders will automatically use the proxy pool—no code changes needed. The middleware will handle rotating proxies, marking dead ones in your database, and retrying failed requests seamlessly.

内容的提问来源于stack exchange,提问作者Paulo Cirino

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:35:39