You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3使用Selenium多进程调用Chrome的问题排查

Hey there, let's work through your two issues step by step—they're both common pitfalls when running Selenium with multiprocessing, so we can get this sorted out smoothly.

First, let's spot a critical bug in your code right away: your ackid_list is defined inside the main() function, but you're trying to use it in the if __name__ == '__main__' block before calling main(). That would throw a NameError because the variable doesn't exist yet. We'll fix that first, then tackle your two main problems.

Problem A: Browser windows won't close

The core issue here is incomplete error handling—if an exception pops up after the initial try/except block (like a download timeout or file rename failure), your browser.quit() call never runs, leaving Chrome processes hanging in the background. Multiprocessing environments can also leave zombie processes if cleanup isn't enforced properly.

Problem B: Shader cache error & program stopping

This error (Failed to create shader cache entry -2) happens because all your Chrome instances are fighting over the same default user data/cache directory. When multiple processes write to the same cache files simultaneously, it causes conflicts that crash Chrome and stop your program.


Fixed Code with Explanations

Here's the revised script with fixes for all issues, plus the initial variable scope bug:

'''Downloading 5500 forms from ERISA'''
# Import Library
from selenium import webdriver
from selenium.webdriver.common.by import By
import time, os, shutil, tempfile
import pandas as pd
from multiprocessing import Pool

# Clean up download folder
DOWNLOAD_DIR = 'D:/Form5500_Downloads'
if os.path.exists(DOWNLOAD_DIR):
    shutil.rmtree(DOWNLOAD_DIR)
os.makedirs(DOWNLOAD_DIR, exist_ok=True)

'''Function to download a single form using ACK ID'''
def download_form(ackid):
    # Create a unique temp directory for each Chrome instance to avoid cache conflicts
    with tempfile.TemporaryDirectory() as temp_profile_dir:
        # Setting Chrome preferences
        chromeOptions = webdriver.ChromeOptions()
        prefs = {"download.default_directory": DOWNLOAD_DIR}
        chromeOptions.add_experimental_option("prefs", prefs)
        
        # Fixes for process cleanup and cache conflicts
        chromeOptions.add_argument(f"--user-data-dir={temp_profile_dir}")  # Unique profile per instance
        chromeOptions.add_argument("--no-sandbox")  # Fixes permission issues in multiprocessing
        chromeOptions.add_argument("--disable-dev-shm-usage")  # Avoids shared memory limits
        chromeOptions.add_argument("--headless=new")  # Optional: Run without visible windows (faster)
        chromeOptions.add_argument("--disable-gpu")  # Disables GPU to avoid shader cache issues
        
        path_to_chromedriver = 'D:/401/401k/chromedriver_2.35.exe'
        browser = None
        try:
            browser = webdriver.Chrome(executable_path=path_to_chromedriver, options=chromeOptions)
            browser.implicitly_wait(10)  # Longer implicit wait to handle slow page loads

            # Open ERISA website
            url = 'https://www.efast.dol.gov/portal/app/disseminate?execution=e1s4#'
            browser.get(url)

            # Search for a form using ACK ID
            browser.find_element(By.CSS_SELECTOR, '#ackId').send_keys(ackid)
            browser.find_element(By.CSS_SELECTOR, '.ui-icon-search').click()

            # Check if the form exists - if NOT, exit the function
            try:
                browser.find_element(By.CSS_SELECTOR, '#form\:filingTreeTable\:0\:einLnk').click()
            except:
                print(f"No form found for ACK ID: {ackid}")
                return  # Exit early if no form exists

            # Wait until downloaded and rename using ackid
            print(f"Processing: {ackid}")
            download_path = os.path.join(DOWNLOAD_DIR, 'filing.pdf')
            timeout = 60  # 1 minute timeout to avoid hanging forever
            start_time = time.time()
            while not os.path.exists(download_path):
                if time.time() - start_time > timeout:
                    print(f"Download timed out for ACK ID: {ackid}")
                    return
                time.sleep(2)  # Shorter sleep interval for responsiveness
            
            # Handle case where file is still being written (Chrome sometimes locks the file)
            time.sleep(1)
            os.rename(download_path, os.path.join(DOWNLOAD_DIR, f"{ackid}.pdf"))
            print(f"Successfully downloaded: {ackid}")

        except Exception as e:
            print(f"Error processing {ackid}: {str(e)}")
        finally:
            # Ensure browser is closed no matter what happens
            if browser:
                browser.quit()
                # Force kill any remaining Chrome processes (extra safety)
                os.system(f"taskkill /im chrome.exe /f /t")

def main():
    # Get list of ackids from csv file
    df = pd.read_csv('D:/401/401k/F_SCH_H_2015_latest.csv', usecols=[0], nrows=10000)
    return df['ACK_ID'].tolist()

if __name__ == '__main__':
    ackid_list = main()
    # Use a smaller pool size if you run into resource issues (10 might be too many for your system)
    with Pool(5) as p:
        p.map(download_form, ackid_list)

Key Fixes Breakdown

  1. Variable Scope Fix: Moved ackid_list retrieval to main() and called it before starting the Pool, so the variable exists when needed.
  2. Unique Chrome Profiles: Used tempfile.TemporaryDirectory() to create a unique user data directory for each Chrome instance—this eliminates the shader cache conflict entirely.
  3. Robust Cleanup: Added a try/finally block to ensure browser.quit() runs no matter what exceptions occur. Also added a taskkill command as a last resort to kill any stuck Chrome processes.
  4. Timeout for Downloads: Added a 1-minute timeout to avoid infinite loops if a download gets stuck.
  5. Chrome Stability Flags: Added --no-sandbox, --disable-dev-shm-usage, and --disable-gpu to prevent crashes and process leaks in multiprocessing environments.
  6. Optional Headless Mode: Enabled --headless=new to run Chrome without visible windows—this uses less system resources and avoids cluttering your screen with dozens of windows. You can remove this line if you need to see the browser for debugging.

Additional Tips

  • Adjust Pool Size: 10 concurrent Chrome instances might overwhelm your system's memory/CPU. Start with 5-6 and increase if your system can handle it.
  • Update ChromeDriver: Your ChromeDriver version (2.35) is very old—make sure it matches your installed Chrome version. Outdated drivers can cause unexpected crashes and compatibility issues.
  • Check Download Locking: Sometimes Chrome keeps the filing.pdf file locked for a second after download—adding a short time.sleep(1) before renaming prevents "file in use" errors.

内容的提问来源于stack exchange,提问作者Ankit Srivastava

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 06:42:06