You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python+Selenium多线程脚本优化:任务切换耗时过长排查

优化Selenium+undetected_chromedriver任务切换耗时问题

问题背景

基于Selenium和undetected_chromedriver开发的自动化脚本,从Excel读取用户名密码列表,多线程处理每个用户的登录、页面PDF生成任务。当前问题:任务完成关闭Driver后,启动新任务耗时约5秒,需要优化缩短该时间。

核心耗时原因

  1. 重复启动浏览器进程:每个任务新建uc.Chrome()实例,Chrome启动本身需要几秒,加上undetected_chromedriver的反检测补丁注入,进一步拉长启动时间
  2. 线程任务管理冗余:手动轮询任务状态的循环增加了不必要的等待开销
  3. 固定等待与无限重试:time.sleep(1.5)的固定等待、异常时的无限循环重试,浪费不必要的时间

具体优化方案

1. 复用浏览器实例(核心优化)

不再为每个用户新建浏览器进程,而是在单个浏览器中用新标签页处理不同用户任务,彻底消除进程启动/销毁的耗时。undetected_chromedriver支持多标签独立操作,每个标签页可登录不同账号。

2. 精简Chrome启动参数

移除重复、非必要参数,减少浏览器启动负载:

  • 移除重复的--window-size配置
  • 关闭useAutomationExtension(原脚本设置为True,会加载冗余自动化扩展)
  • 改用新版无头模式--headless=new,更高效且反检测能力更强

3. 简化线程池逻辑

用concurrent.futures.ThreadPoolExecutor原生方法管理任务,无需手动维护队列和轮询,减少代码复杂度与额外等待。

4. 替换固定等待为显式等待

将time.sleep(1.5)替换为基于元素加载的显式等待,既保证可靠性又节省无意义等待时间。


优化后完整代码

import undetected_chromedriver as uc
import concurrent.futures
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException, TimeoutException
import pandas as pd
import time
import base64
import os

# 精简后的Chrome配置
user_agent = 'Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/15.0 Safari/605.1.15'
chrome_options = webdriver.ChromeOptions()
chrome_options.add_argument('--no-sandbox')
chrome_options.add_argument('--headless=new')  # 新版无头模式,性能与反检测更优
chrome_options.add_argument('--disable-infobars')
chrome_options.add_argument('--disable-extensions')
chrome_options.add_argument('--enable-javascript')
chrome_options.add_argument('--disable-gpu')
chrome_options.add_argument(f'User-Agent={user_agent}')
chrome_options.add_argument('--window-size=2000x2000')
chrome_options.add_argument('--kiosk-printing')
chrome_options.add_argument('--disable-dev-shm-usage')
chrome_options.add_experimental_option("excludeSwitches", ["enable-automation"])
chrome_options.add_experimental_option('useAutomationExtension', False)  # 关闭冗余自动化扩展

# 读取用户列表
df = pd.read_excel("users.xlsx")
username_list = df['username'].tolist()
password_list = df['password'].tolist()

def login_and_generate_pdf(driver, username, password):
    try:
        # 打开新标签页处理当前用户任务
        driver.execute_script("window.open('');")
        driver.switch_to.window(driver.window_handles[-1])
        
        # 登录流程
        driver.get("******")
        WebDriverWait(driver, 5).until(EC.visibility_of_element_located((By.NAME, "username")))
        
        username_field = driver.find_element(By.NAME, "username")
        username_field.send_keys(username)
        
        password_field = driver.find_element(By.NAME, "password")
        password_field.send_keys(password)
        
        submit_btn = WebDriverWait(driver, 2).until(EC.element_to_be_clickable((By.ID, "submitBtn")))
        driver.execute_script("arguments[0].click();", submit_btn)
        
        # 等待登录成功跳转
        WebDriverWait(driver, 3).until(EC.presence_of_element_located((By.XPATH, "//a[text()='s']")))
        
        # 跳转到目标页面并等待加载完成
        driver.get("******")
        WebDriverWait(driver, 3).until(EC.presence_of_element_located((By.TAG_NAME, "body")))
        
        # 生成PDF
        pdf = driver.execute_cdp_cmd(
            "Page.printToPDF", {
                "printBackground": True,
                "landscape": False,
                "displayHeaderFooter": False,
                "scale": 1,
            })
        with open(f'{username}.pdf', "wb") as f:
            f.write(base64.b64decode(pdf['data']))
        
        print(f"用户 {username} 的PDF生成完成")
        
        # 关闭当前标签页,切回主标签页
        driver.close()
        driver.switch_to.window(driver.window_handles[0])
        
    except (NoSuchElementException, ElementClickInterceptedException, TimeoutException) as e:
        print(f"用户 {username} 处理失败: {str(e)}")
        # 异常时清理当前标签页
        if len(driver.window_handles) > 1:
            driver.close()
            driver.switch_to.window(driver.window_handles[0])

def run_tasks():
    max_workers = 2  # 根据机器性能调整线程数
    # 仅启动一次浏览器实例
    with uc.Chrome(options=chrome_options) as driver:
        driver.get("about:blank")  # 初始化主标签页
        
        # 线程池批量处理任务
        with concurrent.futures.ThreadPoolExecutor(max_workers=max_workers) as executor:
            futures = [
                executor.submit(login_and_generate_pdf, driver, username, password) 
                for username, password in zip(username_list, password_list)
            ]
            
            # 等待所有任务完成并捕获异常
            for future in concurrent.futures.as_completed(futures):
                try:
                    future.result()
                except Exception as exc:
                    print(f"任务执行异常: {exc}")

if __name__ == "__main__":
    start_total = time.monotonic()
    run_tasks()
    print(f"所有任务完成,总耗时: {time.monotonic() - start_total:.2f} 秒")

关键改动说明

  • 浏览器复用:仅启动一次Chrome,所有任务通过新标签页处理,完全消除重复启动浏览器的5秒耗时
  • 参数精简:移除冗余配置,关闭自动化扩展,改用新版无头模式,进一步缩短启动时间
  • 线程逻辑简化:用原生方法管理任务,避免手动轮询的额外等待
  • 等待优化:用显式等待替代固定sleep,减少无意义时间浪费
  • 异常处理:异常时及时清理标签页,避免影响后续任务执行

内容的提问来源于stack exchange,提问作者Omar Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 08:56:05