Python concurrent.futures多线程配代理爬取站点任务提前终止如何解决
问题根因
concurrent.futures.Executor.map() 传入多个可迭代对象作为参数时,会以最短的可迭代对象长度作为总执行次数。你的代理列表只有2个元素,因此最多只会执行2次extract调用,任务自然提前终止。同时原生map只会按顺序给每次调用分配参数,无法实现「固定线程绑定固定代理」的需求。
解决方案
方案1:代理轮询(改动最小,适配普通爬取场景)
如果不需要严格绑定线程和代理,仅需要代理重复使用直到所有URL爬完,可以用itertools.cycle把代理列表转为无限循环的迭代器:
import itertools import concurrent.futures # 构造无限循环的代理迭代器 proxy_cycle = itertools.cycle(proxy_list) with concurrent.futures.ThreadPoolExecutor(max_workers=2) as executor: executor.map(extract, url_list, proxy_cycle)
运行时代理会按顺序循环复用,直到所有URL遍历完成。
方案2:严格绑定线程与代理(完全匹配需求)
如果要求每个线程全程固定使用同一个代理,可以用threading.local()为每个线程存储专属代理,每个线程仅在第一次运行时分配一次代理:
- 首先改造
extract函数与代理分配逻辑:
import threading import requests from bs4 import BeautifulSoup thread_local = threading.local() proxy_list = ["1.1.1.1", "2.2.2.2"] # 你的代理列表 def get_thread_proxy(): # 每个线程仅分配一次代理 if not hasattr(thread_local, "proxy"): # 按线程序号对应分配代理 thread_idx = int(threading.current_thread().name.split("-")[-1]) - 1 thread_local.proxy = proxy_list[thread_idx] return thread_local.proxy def extract(url): proxy = get_thread_proxy() print(f'Thread Name : {threading.current_thread().name}') print(f'We are using this proxy : {proxy}') headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:80.0) Gecko/20100101 Firefox/80.0'} try: r = requests.get(url, headers=headers, proxies={'http' : proxy,'https': proxy}, timeout=2) soup = BeautifulSoup(r.text, 'html.parser') page_title = soup.find('title').text.strip() print(page_title) except: pass
- 调用时仅需要传入URL列表即可:
url_list = ["a.com", "b.com", "c.com", "d.com", "e.com"] # 你的站点列表 with concurrent.futures.ThreadPoolExecutor(max_workers=len(proxy_list)) as executor: executor.map(extract, url_list)
注意:该方案需要将
max_workers设置为不超过代理列表的长度,避免索引越界。
内容的提问来源于stack exchange,提问作者shofyankhan
相关产品推荐
相关产品推荐

