多线程爬虫代理轮换代码优化需求及问题咨询
Hey Paul, let's work through optimizing your proxy-rotating multi-threaded crawler. Your core logic makes sense, but we can squash those infinite recursion risks and tighten up proxy error handling with some targeted changes. Here's a step-by-step breakdown and improved implementation:
Key Problem Fixes & Enhancements
1. Kill the Recursive Infinite Loop
Your current make_request uses recursion for retries, which can lead to stack overflow if all proxies fail or the target site keeps blocking requests indefinitely. The fix is to replace recursion with a bounded retry loop that caps the number of attempts:
- Add a
max_retriesparameter (we'll use 5 as a starting point) - Track retry attempts in a loop instead of calling the function recursively
- Exit gracefully when retries are exhausted
2. Strengthen Proxy Error Handling
Your proxy validation and error logic has gaps we can fill:
- The
is_alivecheck uses Google, which might not reflect how the proxy performs on your target site (swap in a lightweight endpoint from your target domain if possible) - Not all
ConnectionErrorinstances mean the proxy is dead—we'll add more granular checks - Ensure proxy removal and selection is fully thread-safe (your lock is a good start, we'll refine it)
- Avoid abrupt program termination with
os._exit(1)—we'll use exceptions or thread-specific exits instead
Bonus Optimizations
- Randomize User-Agent per request instead of once at initialization
- Add jitter to sleep times (random intervals instead of fixed 1s) to mimic human behavior
- Clean up old sessions when switching proxies to avoid stale connections
- Add more robust file loading with error handling for missing
user_agents.txtorproxies.txt
Improved Code Implementation
import requests import threading import random import time import logging from typing import Optional, Dict, Any logging.basicConfig(level=logging.INFO) class Crawler(): def __init__(self): self.user_agents = self._load_user_agents() self.proxies = self._load_proxies() self.lock = threading.Lock() self.current_proxy = None self.counter = 0 self.session = requests.Session() # Initialize with a working proxy self.set_proxy() def _load_user_agents(self) -> list[str]: """Load user agents from file with fallback""" agents = [] try: with open('user_agents.txt', 'r') as f: agents = [line.strip() for line in f if line.strip()] except FileNotFoundError: logging.error("user_agents.txt not found! Using default UA") agents = ["Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36"] return agents def _load_proxies(self) -> list[Dict[str, str]]: """Load proxies from file into request-ready format with fallback""" proxies = [] try: with open('proxies.txt', 'r') as f: for line in f: line = line.strip() if line: proxies.append({ "http": f"http://{line}", "https": f"https://{line}" }) except FileNotFoundError: logging.error("proxies.txt not found! No proxies available") return proxies def _get_random_ua(self) -> str: """Return a random user agent for each request""" return random.choice(self.user_agents) def is_alive(self, proxy: Dict[str, str]) -> bool: """Check if proxy works (replace test URL with your target's lightweight endpoint if possible)""" test_url = "http://www.google.com" try: resp = requests.get(test_url, proxies=proxy, timeout=3, verify=False) return resp.status_code == 200 except Exception: return False def set_proxy(self, remove_current: bool = False) -> None: """Thread-safe proxy selection with validation""" with self.lock: if remove_current and self.current_proxy in self.proxies: try: self.proxies.remove(self.current_proxy) logging.info(f"Removed dead proxy: {self.current_proxy}") except ValueError: pass # Find a working proxy valid_proxy = None while self.proxies and not valid_proxy: candidate = random.choice(self.proxies) if self.is_alive(candidate): valid_proxy = candidate else: self.proxies.remove(candidate) logging.info(f"Removed unresponsive proxy: {candidate}") if not valid_proxy: logging.critical("No working proxies left! Exiting thread.") # Raise exception instead of killing entire program raise RuntimeError("Proxy pool exhausted") # Update session with new proxy and reset counter self.current_proxy = valid_proxy self.session.close() # Clean up old session self.session = requests.Session() self.session.proxies = self.current_proxy self.counter = 0 logging.info(f"Switched to proxy: {self.current_proxy}") def make_request(self, method: str, url: str, max_retries: int = 5, **kwargs) -> Optional[Any]: """Bounded retry loop instead of recursion to avoid infinite loops""" retry_count = 0 while retry_count < max_retries: try: # Rotate proxy after 10-20 requests (thread-safe) with self.lock: if self.counter >= random.randint(10, 20): self.set_proxy() self.counter += 1 # Use fresh UA per request, merge with custom headers if provided headers = {'User-agent': self._get_random_ua()} headers.update(kwargs.pop('headers', {})) if method.upper() == 'GET': if kwargs.pop('download', False): resp = self.session.get(url, headers=headers, stream=True, verify=False, **kwargs) return resp.raw resp = self.session.get(url, headers=headers, verify=False, **kwargs) else: resp = self.session.post(url, headers=headers, verify=False, **kwargs) resp.raise_for_status() # Trigger HTTPError for bad status codes # Handle target-specific block detection html = resp.text if resp.encoding in ['utf8', 'utf-8', None] else resp.content.decode(resp.encoding) if 'Access Denied' in html: logging.warning(f"Proxy blocked by target site: {self.current_proxy}") self.set_proxy(remove_current=True) retry_count += 1 time.sleep(random.uniform(1, 3)) # Random jitter to avoid detection continue return html except requests.exceptions.HTTPError as e: if e.response.status_code in [403, 429]: logging.warning(f"Proxy blocked (HTTP {e.response.status_code}): {self.current_proxy}") self.set_proxy(remove_current=True) elif e.response.status_code == 404: logging.error(f"URL not found: {url}") return None else: logging.error(f"HTTP Error for {url}: {str(e)}") retry_count += 1 time.sleep(random.uniform(1, 3)) except requests.exceptions.Timeout: logging.warning(f"Request timed out for {url}") retry_count += 1 time.sleep(random.uniform(2, 4)) except requests.exceptions.ConnectionError as e: if "403 Forbidden" in str(e): logging.critical("Global 403 block detected—check your IP or user agents") return None logging.warning(f"Connection error with proxy {self.current_proxy}: {str(e)}") self.set_proxy(remove_current=True) retry_count += 1 time.sleep(random.uniform(2, 4)) except RuntimeError as e: logging.error(str(e)) return None except Exception as e: logging.error(f"Unexpected error for {url}: {str(e)}") retry_count += 1 time.sleep(random.uniform(1, 3)) logging.error(f"Max retries ({max_retries}) exhausted for {url}") return None
Key Changes Explained
- Bounded Retries: Replaced recursion with a
whileloop that stops aftermax_retriesattempts, eliminating infinite recursion risks - Thread-Safe Proxy Management: Locked all proxy pool modifications and selection to avoid race conditions in multi-threaded environments
- Granular Error Handling: Split out different error types (timeout, connection issues, HTTP blocks) with appropriate actions
- Randomized Behavior: Per-request User-Agent and jittered sleep times reduce detection risk
- Graceful Exit: Removed
os._exit(1)in favor of exceptions and thread-specific exits, so one failing thread doesn't take down the whole program - Robust File Loading: Added fallback logic for missing
user_agents.txtorproxies.txtfiles
内容的提问来源于stack exchange,提问作者Paul R.
相关产品推荐
相关产品推荐

