You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium+ChromeDriver(Python):如何判断下载完成替代静态等待

Replace Static Sleep with Chrome DevTools Download Monitoring

Great question—static sleeps are definitely a fragile solution, especially with variable download speeds and multi-threaded workflows. Since you can't rely on file system checks (dynamic filenames), we can leverage Chrome's DevTools Protocol (CDP) via Selenium to directly monitor the browser's download manager. Here's a robust, event-driven approach to replace time.sleep():

Step 1: Configure Chrome for Unattended Downloads

First, set up your ChromeDriver to bypass download prompts and use a dedicated download directory. This ensures downloads start automatically without user interaction:

from selenium.webdriver.chrome.options import Options
from selenium import webdriver

def create_thread_safe_driver(download_dir):
    chrome_options = Options()
    download_prefs = {
        "download.default_directory": download_dir,
        "download.prompt_for_download": False,  # Disable download confirmation popups
        "download.directory_upgrade": True,
        "plugins.always_open_pdf_externally": True  # Force PDF download instead of opening in browser
    }
    chrome_options.add_experimental_option("prefs", download_prefs)
    # Optional: Add headless mode if you don't need a visible browser
    # chrome_options.add_argument("--headless=new")
    return webdriver.Chrome(options=chrome_options)

Step 2: Build a Download Waiter Using CDP

We'll use Chrome's CDP to enable download monitoring, track new download items, and wait until they complete. This works per-driver instance, so it's safe for multi-threading:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.common.exceptions import TimeoutException

def wait_for_download(driver, timeout=30):
    # Enable CDP domains for downloads and page monitoring
    driver.execute_cdp_cmd("Page.enable", {})
    driver.execute_cdp_cmd("Downloads.enable", {})

    # Get initial list of downloads to avoid tracking old items
    initial_downloads = driver.execute_cdp_cmd("Downloads.getDownloads", {})["downloads"]
    initial_count = len(initial_downloads)

    def is_download_complete(_driver):
        # Fetch current download state
        current_downloads = _driver.execute_cdp_cmd("Downloads.getDownloads", {})["downloads"]
        # Check only new downloads initiated after clicking the button
        for download in current_downloads[initial_count:]:
            if download["state"] == "completed":
                return True
            elif download["state"] == "interrupted":
                raise RuntimeError(f"Download failed: {download['error']['description']}")
        return False

    try:
        # Wait until the download completes or times out
        WebDriverWait(driver, timeout).until(is_download_complete)
        return True
    finally:
        # Clean up CDP resources regardless of outcome
        driver.execute_cdp_cmd("Downloads.disable", {})

Step 3: Update Your Download Function

Replace the time.sleep(10) with our wait_for_download function. We'll also add better retry logic:

from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
import threading

# Make sure Counter is thread-safe with a lock
Counter = 0
counter_lock = threading.Lock()

def file_download(num, drivervar, download_dir):
    global Counter
    with counter_lock:
        Counter += 1
    try:
        drivervar.get(url[num])
        # Wait for download button to be clickable
        download_button = WebDriverWait(drivervar, 20).until(
            EC.element_to_be_clickable((By.ID, 'download button ID'))
        )
        download_button.click()
        
        # Wait for download to finish instead of static sleep
        if not wait_for_download(drivervar, timeout=30):
            print(f"Thread {num}: Download timed out after 30 seconds")
    except TimeoutException:
        print(f'Timeout in thread number: {num}, retrying...')
        # Retry logic with shorter timeout
        try:
            drivervar.get(url[num])
            download_button = WebDriverWait(drivervar, 10).until(
                EC.element_to_be_clickable((By.ID, 'download button ID'))
            )
            download_button.click()
            wait_for_download(drivervar, timeout=30)
        except TimeoutException:
            print(f'Thread {num}: Retry failed')

Key Advantages Over time.sleep()

  • Event-driven: Proceeds immediately once the download finishes, no wasted time waiting for arbitrary delays.
  • Robust: Catches failed downloads (interrupted state) instead of silently continuing.
  • Thread-safe: Each driver instance monitors its own downloads, so multi-threaded workflows won't interfere with each other.

Important Notes

  • Ensure your ChromeDriver version matches your installed Chrome version (critical for CDP compatibility).
  • If using headless mode, use the --headless=new flag (the old --headless flag has limited download support).
  • For very large files, adjust the timeout parameter in wait_for_download to match your expected maximum download time.

内容的提问来源于stack exchange,提问作者BlackMamba

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:16:26