Python+Selenium多线程脚本优化:任务切换耗时过长排查
优化Selenium+undetected_chromedriver任务切换耗时问题
问题背景
基于Selenium和undetected_chromedriver开发的自动化脚本,从Excel读取用户名密码列表,多线程处理每个用户的登录、页面PDF生成任务。当前问题:任务完成关闭Driver后,启动新任务耗时约5秒,需要优化缩短该时间。
核心耗时原因
- 重复启动浏览器进程:每个任务新建
uc.Chrome()实例,Chrome启动本身需要几秒,加上undetected_chromedriver的反检测补丁注入,进一步拉长启动时间 - 线程任务管理冗余:手动轮询任务状态的循环增加了不必要的等待开销
- 固定等待与无限重试:
time.sleep(1.5)的固定等待、异常时的无限循环重试,浪费不必要的时间
具体优化方案
1. 复用浏览器实例(核心优化)
不再为每个用户新建浏览器进程,而是在单个浏览器中用新标签页处理不同用户任务,彻底消除进程启动/销毁的耗时。undetected_chromedriver支持多标签独立操作,每个标签页可登录不同账号。
2. 精简Chrome启动参数
移除重复、非必要参数,减少浏览器启动负载:
- 移除重复的
--window-size配置 - 关闭
useAutomationExtension(原脚本设置为True,会加载冗余自动化扩展) - 改用新版无头模式
--headless=new,更高效且反检测能力更强
3. 简化线程池逻辑
用concurrent.futures.ThreadPoolExecutor原生方法管理任务,无需手动维护队列和轮询,减少代码复杂度与额外等待。
4. 替换固定等待为显式等待
将time.sleep(1.5)替换为基于元素加载的显式等待,既保证可靠性又节省无意义等待时间。
优化后完整代码
import undetected_chromedriver as uc import concurrent.futures from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException, TimeoutException import pandas as pd import time import base64 import os # 精简后的Chrome配置 user_agent = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/15.0 Safari/605.1.15' chrome_options = webdriver.ChromeOptions() chrome_options.add_argument('--no-sandbox') chrome_options.add_argument('--headless=new') # 新版无头模式,性能与反检测更优 chrome_options.add_argument('--disable-infobars') chrome_options.add_argument('--disable-extensions') chrome_options.add_argument('--enable-javascript') chrome_options.add_argument('--disable-gpu') chrome_options.add_argument(f'User-Agent={user_agent}') chrome_options.add_argument('--window-size=2000x2000') chrome_options.add_argument('--kiosk-printing') chrome_options.add_argument('--disable-dev-shm-usage') chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"]) chrome_options.add_experimental_option('useAutomationExtension', False) # 关闭冗余自动化扩展 # 读取用户列表 df = pd.read_excel("users.xlsx") username_list = df['username'].tolist() password_list = df['password'].tolist() def login_and_generate_pdf(driver, username, password): try: # 打开新标签页处理当前用户任务 driver.execute_script("window.open('');") driver.switch_to.window(driver.window_handles[-1]) # 登录流程 driver.get("******") WebDriverWait(driver, 5).until(EC.visibility_of_element_located((By.NAME, "username"))) username_field = driver.find_element(By.NAME, "username") username_field.send_keys(username) password_field = driver.find_element(By.NAME, "password") password_field.send_keys(password) submit_btn = WebDriverWait(driver, 2).until(EC.element_to_be_clickable((By.ID, "submitBtn"))) driver.execute_script("arguments[0].click();", submit_btn) # 等待登录成功跳转 WebDriverWait(driver, 3).until(EC.presence_of_element_located((By.XPATH, "//a[text()='s']"))) # 跳转到目标页面并等待加载完成 driver.get("******") WebDriverWait(driver, 3).until(EC.presence_of_element_located((By.TAG_NAME, "body"))) # 生成PDF pdf = driver.execute_cdp_cmd( "Page.printToPDF", { "printBackground": True, "landscape": False, "displayHeaderFooter": False, "scale": 1, }) with open(f'{username}.pdf', "wb") as f: f.write(base64.b64decode(pdf['data'])) print(f"用户 {username} 的PDF生成完成") # 关闭当前标签页,切回主标签页 driver.close() driver.switch_to.window(driver.window_handles[0]) except (NoSuchElementException, ElementClickInterceptedException, TimeoutException) as e: print(f"用户 {username} 处理失败: {str(e)}") # 异常时清理当前标签页 if len(driver.window_handles) > 1: driver.close() driver.switch_to.window(driver.window_handles[0]) def run_tasks(): max_workers = 2 # 根据机器性能调整线程数 # 仅启动一次浏览器实例 with uc.Chrome(options=chrome_options) as driver: driver.get("about:blank") # 初始化主标签页 # 线程池批量处理任务 with concurrent.futures.ThreadPoolExecutor(max_workers=max_workers) as executor: futures = [ executor.submit(login_and_generate_pdf, driver, username, password) for username, password in zip(username_list, password_list) ] # 等待所有任务完成并捕获异常 for future in concurrent.futures.as_completed(futures): try: future.result() except Exception as exc: print(f"任务执行异常: {exc}") if __name__ == "__main__": start_total = time.monotonic() run_tasks() print(f"所有任务完成,总耗时: {time.monotonic() - start_total:.2f} 秒")
关键改动说明
- 浏览器复用:仅启动一次Chrome,所有任务通过新标签页处理,完全消除重复启动浏览器的5秒耗时
- 参数精简:移除冗余配置,关闭自动化扩展,改用新版无头模式,进一步缩短启动时间
- 线程逻辑简化:用原生方法管理任务,避免手动轮询的额外等待
- 等待优化:用显式等待替代固定sleep,减少无意义时间浪费
- 异常处理:异常时及时清理标签页,避免影响后续任务执行
内容的提问来源于stack exchange,提问作者Omar Ahmed
相关产品推荐
相关产品推荐

