You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Tor代理的多线程爬虫SOCKS报错问题及优化咨询

解决Tor多线程爬虫的SOCKS服务器失败问题及性能优化

咱们先拆解下你遇到的问题:Socket connection failed (Socket error: 0x01: General SOCKS server failure)在多线程下出现、多进程正常,核心原因有两个:

  1. 你当前全局替换socket.socket的方式,会导致所有线程共享同一个Tor SOCKS连接上下文,多线程并发请求时Tor服务器扛不住这么多同时连接,直接返回通用失败;
  2. 200个线程的并发量远超Tor默认的连接限制,相当于短时间内给Tor塞了太多请求,直接触发了拒绝机制。

下面给你一步步的修复和优化方案:

一、修复SOCKS错误:取消全局Socket替换,给每个请求单独配置代理

全局替换socket是多线程爬虫的大忌——所有线程会共用同一个socket上下文,导致连接混乱。换成给每个requests Session单独设置Tor代理,让每个线程的请求都有独立的连接通道:

import requests
from bs4 import BeautifulSoup
import random
from stem import Signal
from stem.control import Controller
import time

# 假设BROWSERS是你的UA列表
BROWSERS = ["Mozilla/5.0 ...", "..."]

def renew_tor():
    # 每次刷新节点都重新创建Controller,避免多线程操作同一实例的线程安全问题
    with Controller.from_port(port=9151) as controller:
        controller.authenticate()
        controller.signal(Signal.NEWNYM)
    # Tor切换节点需要时间,官方建议至少等10秒,避免请求还没切换成功就发出去
    time.sleep(10)
    # 直接返回新的headers,不用全局变量
    return {
        "Accept-Language": "en-US,en;q=0.5",
        "User-Agent": random.choice(BROWSERS),
        "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
        "Referer": "http://thewebsite2.com",
        "Connection": "close"
    }

def get_soup(url):
    while True:
        try:
            # 每个请求创建独立的Session,单独配置Tor代理
            session = requests.Session()
            session.proxies = {
                "http": "socks5://127.0.0.1:9150",
                "https": "socks5://127.0.0.1:9150"
            }
            # 首次请求用随机UA,遇到验证码后刷新Tor并更新headers
            request_headers = renew_tor() if "captcha" in locals() else {
                "Accept-Language": "en-US,en;q=0.5",
                "User-Agent": random.choice(BROWSERS),
                "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
                "Referer": "http://thewebsite2.com",
                "Connection": "close"
            }
            session.headers.update(request_headers)
            
            response = session.get(url, timeout=15)
            response.raise_for_status()  # 主动抛出HTTP错误,比如403、500
            the_page = response.content.decode('utf-8', errors='ignore')
            the_soup = BeautifulSoup(the_page, 'html.parser')
            
            if "captcha" in the_page.lower():
                print(f"Captcha detected for URL: {url}")
                request_headers = renew_tor()
            else:
                return the_soup
        except Exception as e:
            print(f"Error fetching {url}: {str(e)}")
            # 出错后刷新Tor节点,再重试
            renew_tor()
            time.sleep(3)

二、调整并发数,避免Tor过载

Tor的默认并发连接数有限(一般在10-50之间),你开200个线程等于直接把Tor压垮了。把线程池大小降到合理范围,同时给每个请求加随机延迟,避免请求过于集中:

from concurrent import futures
import random
import time

# 把线程池大小调整为30(可根据Tor的负载情况微调)
with futures.ThreadPoolExecutor(30) as executor:
    for url in zurls:
        # 加0.5-2秒的随机延迟,模拟人类浏览行为,减少Tor和目标网站的压力
        time.sleep(random.uniform(0.5, 2))
        executor.submit(fetchjob, url)

三、性能优化建议

1. 复用线程专属的Session

每次请求新建Session会增加TCP连接的开销,用线程局部存储给每个线程分配一个专属Session,复用连接:

import threading

# 线程局部存储,每个线程有自己的Session实例
thread_local = threading.local()

def get_thread_session():
    if not hasattr(thread_local, "session"):
        session = requests.Session()
        session.proxies = {
            "http": "socks5://127.0.0.1:9150",
            "https": "socks5://127.0.0.1:9150"
        }
        thread_local.session = session
    return thread_local.session

# 修改get_soup函数,用线程专属Session
def get_soup(url):
    session = get_thread_session()
    # 每次请求更新UA,避免被识别为爬虫
    session.headers.update({
        "User-Agent": random.choice(BROWSERS),
        # 其他headers保持不变
    })
    # 后续请求逻辑...

2. 用重试库优化异常处理

自己写while True重试不够优雅,用tenacity库实现指数退避重试,减少无效的重复请求:

from tenacity import retry, stop_after_attempt, wait_exponential

@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=10))
def get_soup(url):
    session = get_thread_session()
    session.headers.update({"User-Agent": random.choice(BROWSERS)})
    response = session.get(url, timeout=15)
    response.raise_for_status()
    the_page = response.content.decode('utf-8', errors='ignore')
    the_soup = BeautifulSoup(the_page, 'html.parser')
    
    if "captcha" in the_page.lower():
        print(f"Captcha detected for URL: {url}")
        renew_tor()
        # 触发重试
        raise Exception("Captcha encountered, retrying after Tor renewal")
    return the_soup

3. 监控Tor连接状态

定期检查Tor是否正常运行,避免无效请求浪费资源:

def check_tor_status():
    try:
        session = requests.Session()
        session.proxies = {"http": "socks5://127.0.0.1:9150", "https": "socks5://127.0.0.1:9150"}
        response = session.get("https://check.torproject.org/", timeout=10)
        return "Congratulations. This browser is configured to use Tor." in response.text
    except Exception as e:
        print(f"Tor connection failed: {str(e)}")
        return False

# 在爬虫启动前检查Tor状态
if not check_tor_status():
    print("Tor is not running properly, exiting...")
    exit(1)

内容的提问来源于stack exchange,提问作者user5236897

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 03:41:36