如何在Python Selenium爬虫中使用多进程调用函数并避免多浏览器驱动冲突
如何在Python Selenium爬虫中使用多进程调用函数并避免多浏览器驱动冲突
你遇到的问题太典型了——多进程下共享同一个Selenium驱动实例,必然会导致各种混乱冲突,毕竟多个进程同时操作同一个浏览器窗口,结果肯定是互相干扰。解决的核心思路就是给每个进程分配完全独立的浏览器驱动实例,让它们各自跑自己的任务,彻底隔离资源。
核心修改思路
- 砍掉全局的
driver和data变量,让每个进程自己初始化专属的浏览器实例和数据字典 - 重构所有操作浏览器的函数,让它们接收
driver和data作为参数,不再依赖全局变量 - 用Python的
multiprocessing模块管理进程,每个进程独立执行完整的任务流程 - 确保每个进程结束后正确关闭浏览器驱动,避免残留僵尸进程
修改后的完整代码示例
# importing important libraries from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.common.action_chains import ActionChains import time from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.chrome.options import Options from test_sheet import connection_sheet import multiprocessing # 封装浏览器配置,每个进程都能创建独立的配置 def get_browser_options(): options = Options() # options.add_experimental_option("detach", True) # options.add_argument('headless') # 开启无头模式可节省资源,需要时取消注释 return options # 重构爬虫核心函数,接收driver和data参数,不再依赖全局变量 def find_and_fetch(driver, data): try: # Use JavaScript to hide the banner driver.execute_script(""" var banner = document.querySelector('a.message-banner'); if (banner) { banner.style.display = 'none'; }""") except: print(f"banner disabled (进程: {multiprocessing.current_process().name})") time.sleep(1) try: current_title = driver.find_element(By.XPATH,"//h5[@class='ng-binding']").text if not current_title: raise Exception("No title found") # 处理器信息爬取逻辑 if 'Processor' in current_title: processor = [] try: elem = driver.find_element(By.XPATH,"//div[@class='answers']") elems = elem.find_elements(By.XPATH, "//div[@ng-keydown='selectAnswer($index, $event)']") for i in elems: processor.append(i.get_attribute('aria-label')) for prces in processor: time.sleep(1) data['Processor']=prces print(f"=> processor : {prces} (进程: {multiprocessing.current_process().name})") find_and_click(driver, prces) find_and_fetch(driver, data) except Exception as e : print(f"Error While passing to processor selection! {e}") data['Processor']='-' return # 内存容量爬取逻辑 if "memory capacity" in current_title: capas = [] try: elem = driver.find_element(By.XPATH,"//div[@class='answers']") elems = elem.find_elements(By.XPATH, "//div[@ng-keydown='selectAnswer($index, $event)']") for i in elems: capas.append(i.get_attribute('aria-label')) for capa in capas: time.sleep(1) data['Memory Capacity']=capa print(f"=> memory capacity : {capa} (进程: {multiprocessing.current_process().name})") find_and_click(driver, capa) find_and_fetch(driver, data) small_back_button(driver) except : print(f"Error While passing to memory capacity! (进程: {multiprocessing.current_process().name})") data['Memory Capacity']='-' return # 存储容量爬取逻辑 if "storage capacity" in current_title: storage = [] try: elem = driver.find_element(By.XPATH,"//div[@class='answers']") elems = elem.find_elements(By.XPATH, "//div[@ng-keydown='selectAnswer($index, $event)']") for i in elems: storage.append(i.get_attribute('aria-label')) for store in storage: time.sleep(1) start_time = time.time() data['Storage Capacity']=store print(f"=> storage capacity : {store} (进程: {multiprocessing.current_process().name})") find_and_click(driver, store) find_and_fetch(driver, data) end_time = time.time() time_taken = end_time-start_time print(f"The code block took {time_taken:.4f} seconds to execute.") driver.quit() break small_back_button(driver) except : print(f"Error While passing to storage selection! (进程: {multiprocessing.current_process().name})") data['Storage Capacity']='-' return # 以下Condition、Battery Health、Include Charger等逻辑 # 请按照上面的格式,把driver和data作为参数传入修改原代码对应部分 # 这里省略重复代码块,你可以自行补充 except: pass # 最终价格页处理逻辑 time.sleep(1) try: offer_text = driver.find_element(By.XPATH, "//h3[@class='your-offer']").text if "Your device is valued at" in offer_text: print(f'in the final page (进程: {multiprocessing.current_process().name})') fetch_info(driver, data) time.sleep(1) large_back_button(driver) return except: pass # 无法匹配时回退 try: small_back_button(driver) except: pass return # 重构点击函数,接收driver参数 def find_and_click(driver, elem_text): time.sleep(1) try: elem = WebDriverWait(driver, 5).until( EC.element_to_be_clickable((By.XPATH, f"//div[@aria-label='{elem_text}']")) ) driver.execute_script("arguments[0].scrollIntoView(true);", elem) action = ActionChains(driver) action.move_to_element(elem).click().perform() except: print(f"could not find any clickable element: {elem_text} (进程: {multiprocessing.current_process().name})") # 以下辅助函数都需要重构,接收driver和data参数 def small_back_button(driver): # 补充你原有的small_back_button逻辑,比如: try: driver.find_element(By.XPATH, "//你的返回按钮XPATH").click() except: print(f"small back button not found (进程: {multiprocessing.current_process().name})") def large_back_button(driver): # 补充你原有的large_back_button逻辑 pass def fetch_info(driver, data): # 补充你原有的fetch_info逻辑,比如爬取价格并存入data try: price = driver.find_element(By.XPATH, "//价格元素XPATH").text data['Price'] = price # 注意:如果要写入sheet,需要加进程锁避免冲突 # with lock: # connection_sheet.write(data) except: print(f"failed to fetch info (进程: {multiprocessing.current_process().name})") # 每个进程的完整任务流程 def process_task(): # 每个进程初始化自己的浏览器、配置和数据 options = get_browser_options() driver = webdriver.Firefox(options=options) data = { 'Device':'-', 'Processor':'-', 'Memory Capacity':'-', 'Storage Capacity':'-', 'Condition':'-', 'Battery Health':'-', 'Include Charger':'-', 'Fully Functional':'-', 'Price':'-' } # 打开目标网页(补充你的目标URL) # driver.get("https://your-target-url.com") # 执行爬虫任务 find_and_fetch(driver, data) # 任务结束后关闭驱动 driver.quit() if __name__ == '__main__': # 进程数量根据机器性能调整,不要开太多 process_count = 2 processes = [] # 如果需要多进程写数据,初始化锁 # lock = multiprocessing.Lock() for i in range(process_count): p = multiprocessing.Process(target=process_task, name=f"Task-{i+1}") processes.append(p) p.start() # 等待所有进程完成 for p in processes: p.join()
额外注意事项
- 多进程数据写入安全:如果你的
connection_sheet涉及多进程同时写操作,一定要用multiprocessing.Lock做同步,否则会导致数据覆盖或损坏。 - 资源控制:不要盲目开大量进程,每个进程对应一个浏览器窗口,过多会占用大量内存CPU,反而降低效率,建议根据机器性能设置2-4个进程。
- 驱动版本匹配:确保Firefox浏览器和geckodriver版本完全匹配,多进程下版本不兼容问题会被放大。
- 无头模式:如果不需要可视化浏览器,开启
headless模式能大幅节省系统资源,修改get_browser_options函数即可。
备注:内容来源于stack exchange,提问作者Deep
相关产品推荐
相关产品推荐

