You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

多线程爬虫代理轮换代码优化需求及问题咨询

Hey Paul, let's work through optimizing your proxy-rotating multi-threaded crawler. Your core logic makes sense, but we can squash those infinite recursion risks and tighten up proxy error handling with some targeted changes. Here's a step-by-step breakdown and improved implementation:

Key Problem Fixes & Enhancements

1. Kill the Recursive Infinite Loop

Your current make_request uses recursion for retries, which can lead to stack overflow if all proxies fail or the target site keeps blocking requests indefinitely. The fix is to replace recursion with a bounded retry loop that caps the number of attempts:

  • Add a max_retries parameter (we'll use 5 as a starting point)
  • Track retry attempts in a loop instead of calling the function recursively
  • Exit gracefully when retries are exhausted

2. Strengthen Proxy Error Handling

Your proxy validation and error logic has gaps we can fill:

  • The is_alive check uses Google, which might not reflect how the proxy performs on your target site (swap in a lightweight endpoint from your target domain if possible)
  • Not all ConnectionError instances mean the proxy is dead—we'll add more granular checks
  • Ensure proxy removal and selection is fully thread-safe (your lock is a good start, we'll refine it)
  • Avoid abrupt program termination with os._exit(1)—we'll use exceptions or thread-specific exits instead

Bonus Optimizations

  • Randomize User-Agent per request instead of once at initialization
  • Add jitter to sleep times (random intervals instead of fixed 1s) to mimic human behavior
  • Clean up old sessions when switching proxies to avoid stale connections
  • Add more robust file loading with error handling for missing user_agents.txt or proxies.txt

Improved Code Implementation
import requests
import threading
import random
import time
import logging
from typing import Optional, Dict, Any

logging.basicConfig(level=logging.INFO)

class Crawler():
    def __init__(self):
        self.user_agents = self._load_user_agents()
        self.proxies = self._load_proxies()
        self.lock = threading.Lock()
        self.current_proxy = None
        self.counter = 0
        self.session = requests.Session()
        # Initialize with a working proxy
        self.set_proxy()

    def _load_user_agents(self) -> list[str]:
        """Load user agents from file with fallback"""
        agents = []
        try:
            with open('user_agents.txt', 'r') as f:
                agents = [line.strip() for line in f if line.strip()]
        except FileNotFoundError:
            logging.error("user_agents.txt not found! Using default UA")
            agents = ["Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"]
        return agents

    def _load_proxies(self) -> list[Dict[str, str]]:
        """Load proxies from file into request-ready format with fallback"""
        proxies = []
        try:
            with open('proxies.txt', 'r') as f:
                for line in f:
                    line = line.strip()
                    if line:
                        proxies.append({
                            "http": f"http://{line}",
                            "https": f"https://{line}"
                        })
        except FileNotFoundError:
            logging.error("proxies.txt not found! No proxies available")
        return proxies

    def _get_random_ua(self) -> str:
        """Return a random user agent for each request"""
        return random.choice(self.user_agents)

    def is_alive(self, proxy: Dict[str, str]) -> bool:
        """Check if proxy works (replace test URL with your target's lightweight endpoint if possible)"""
        test_url = "http://www.google.com"
        try:
            resp = requests.get(test_url, proxies=proxy, timeout=3, verify=False)
            return resp.status_code == 200
        except Exception:
            return False

    def set_proxy(self, remove_current: bool = False) -> None:
        """Thread-safe proxy selection with validation"""
        with self.lock:
            if remove_current and self.current_proxy in self.proxies:
                try:
                    self.proxies.remove(self.current_proxy)
                    logging.info(f"Removed dead proxy: {self.current_proxy}")
                except ValueError:
                    pass

            # Find a working proxy
            valid_proxy = None
            while self.proxies and not valid_proxy:
                candidate = random.choice(self.proxies)
                if self.is_alive(candidate):
                    valid_proxy = candidate
                else:
                    self.proxies.remove(candidate)
                    logging.info(f"Removed unresponsive proxy: {candidate}")

            if not valid_proxy:
                logging.critical("No working proxies left! Exiting thread.")
                # Raise exception instead of killing entire program
                raise RuntimeError("Proxy pool exhausted")

            # Update session with new proxy and reset counter
            self.current_proxy = valid_proxy
            self.session.close()  # Clean up old session
            self.session = requests.Session()
            self.session.proxies = self.current_proxy
            self.counter = 0
            logging.info(f"Switched to proxy: {self.current_proxy}")

    def make_request(self, method: str, url: str, max_retries: int = 5, **kwargs) -> Optional[Any]:
        """Bounded retry loop instead of recursion to avoid infinite loops"""
        retry_count = 0
        while retry_count < max_retries:
            try:
                # Rotate proxy after 10-20 requests (thread-safe)
                with self.lock:
                    if self.counter >= random.randint(10, 20):
                        self.set_proxy()
                    self.counter += 1

                # Use fresh UA per request, merge with custom headers if provided
                headers = {'User-agent': self._get_random_ua()}
                headers.update(kwargs.pop('headers', {}))

                if method.upper() == 'GET':
                    if kwargs.pop('download', False):
                        resp = self.session.get(url, headers=headers, stream=True, verify=False, **kwargs)
                        return resp.raw
                    resp = self.session.get(url, headers=headers, verify=False, **kwargs)
                else:
                    resp = self.session.post(url, headers=headers, verify=False, **kwargs)

                resp.raise_for_status()  # Trigger HTTPError for bad status codes

                # Handle target-specific block detection
                html = resp.text if resp.encoding in ['utf8', 'utf-8', None] else resp.content.decode(resp.encoding)
                if 'Access Denied' in html:
                    logging.warning(f"Proxy blocked by target site: {self.current_proxy}")
                    self.set_proxy(remove_current=True)
                    retry_count += 1
                    time.sleep(random.uniform(1, 3))  # Random jitter to avoid detection
                    continue

                return html

            except requests.exceptions.HTTPError as e:
                if e.response.status_code in [403, 429]:
                    logging.warning(f"Proxy blocked (HTTP {e.response.status_code}): {self.current_proxy}")
                    self.set_proxy(remove_current=True)
                elif e.response.status_code == 404:
                    logging.error(f"URL not found: {url}")
                    return None
                else:
                    logging.error(f"HTTP Error for {url}: {str(e)}")

                retry_count += 1
                time.sleep(random.uniform(1, 3))

            except requests.exceptions.Timeout:
                logging.warning(f"Request timed out for {url}")
                retry_count += 1
                time.sleep(random.uniform(2, 4))

            except requests.exceptions.ConnectionError as e:
                if "403 Forbidden" in str(e):
                    logging.critical("Global 403 block detected—check your IP or user agents")
                    return None
                logging.warning(f"Connection error with proxy {self.current_proxy}: {str(e)}")
                self.set_proxy(remove_current=True)
                retry_count += 1
                time.sleep(random.uniform(2, 4))

            except RuntimeError as e:
                logging.error(str(e))
                return None

            except Exception as e:
                logging.error(f"Unexpected error for {url}: {str(e)}")
                retry_count += 1
                time.sleep(random.uniform(1, 3))

        logging.error(f"Max retries ({max_retries}) exhausted for {url}")
        return None

Key Changes Explained
  • Bounded Retries: Replaced recursion with a while loop that stops after max_retries attempts, eliminating infinite recursion risks
  • Thread-Safe Proxy Management: Locked all proxy pool modifications and selection to avoid race conditions in multi-threaded environments
  • Granular Error Handling: Split out different error types (timeout, connection issues, HTTP blocks) with appropriate actions
  • Randomized Behavior: Per-request User-Agent and jittered sleep times reduce detection risk
  • Graceful Exit: Removed os._exit(1) in favor of exceptions and thread-specific exits, so one failing thread doesn't take down the whole program
  • Robust File Loading: Added fallback logic for missing user_agents.txt or proxies.txt files

内容的提问来源于stack exchange,提问作者Paul R.

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:50:25