如何避免Pool执行非预期的多次重复操作?多进程爬虫场景问题咨询
问题根源
该现象是Windows系统下Python多进程的spawn启动模式导致的:
- Windows没有fork系统调用,子进程启动时会重新导入当前运行的整个py文件作为模块
- 你将webdriver初始化、网页访问的代码直接写在了全局作用域,没有做进程保护,所以每启动一个子进程,都会重新执行一次全局的webdriver初始化逻辑,弹出新的浏览器窗口,和父进程是否提前关闭了driver无关。
解决方案
1. 添加入口保护
把所有全局执行的业务逻辑全部放到if __name__ == '__main__'代码块内,这样子进程导入模块时不会执行这部分逻辑:
import time from selenium import webdriver import concurrent.futures # 函数定义、全局常量可以放在块外 def extraccionFutbolBetfair(url): # 你的原有爬取逻辑 pass if __name__ == '__main__': deporteElegido = 'Fútbol' web = 'https://sports.williamhill.es/betting/es-es' PATH = 'C:\WebDrivers\ChromeDriver\chromedriver.exe' driver = webdriver.Chrome(PATH) inicio = time.time() driver.get('https://www.betfair.es/sport/') driver.implicitly_wait(10) try: driver.find_element_by_id('onetrust-accept-btn-handler').click() except: pass driver.implicitly_wait(10) driver.find_element_by_xpath('//a[@class="ui-expandable allSportsLink allSportsMainLink ui-betslip-action"]').click() driver.implicitly_wait(10) enlaces = driver.find_elements_by_xpath('//li[@class="full-sports-item"]') deportes = [enlace.text for enlace in enlaces] indice = deportes.index(deporteElegido) enlaces[indice].click() time.sleep(2) links = driver.find_elements_by_xpath('//div[@class="avb-col avb-col-runners"]//a') enlaces = [link.get_attribute('href') for link in links] driver.quit() with concurrent.futures.ProcessPoolExecutor() as executor: print([x for x in executor.map(extraccionFutbolBetfair, enlaces)]) fin = time.time() print(fin-inicio)
2. 可选:隐藏浏览器窗口
如果extraccionFutbolBetfair函数本身需要为每个链接初始化webdriver执行爬取,此时弹出窗口属于正常业务逻辑,你可以添加无头模式参数隐藏窗口:
# 在初始化webdriver前添加配置 options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(PATH, options=options)
内容的提问来源于stack exchange,提问作者Juan José Campos
相关产品推荐
相关产品推荐

