Python3使用Selenium多进程调用Chrome的问题排查
Hey there, let's work through your two issues step by step—they're both common pitfalls when running Selenium with multiprocessing, so we can get this sorted out smoothly.
First, let's spot a critical bug in your code right away: your ackid_list is defined inside the main() function, but you're trying to use it in the if __name__ == '__main__' block before calling main(). That would throw a NameError because the variable doesn't exist yet. We'll fix that first, then tackle your two main problems.
Problem A: Browser windows won't close
The core issue here is incomplete error handling—if an exception pops up after the initial try/except block (like a download timeout or file rename failure), your browser.quit() call never runs, leaving Chrome processes hanging in the background. Multiprocessing environments can also leave zombie processes if cleanup isn't enforced properly.
Problem B: Shader cache error & program stopping
This error (Failed to create shader cache entry -2) happens because all your Chrome instances are fighting over the same default user data/cache directory. When multiple processes write to the same cache files simultaneously, it causes conflicts that crash Chrome and stop your program.
Fixed Code with Explanations
Here's the revised script with fixes for all issues, plus the initial variable scope bug:
'''Downloading 5500 forms from ERISA''' # Import Library from selenium import webdriver from selenium.webdriver.common.by import By import time, os, shutil, tempfile import pandas as pd from multiprocessing import Pool # Clean up download folder DOWNLOAD_DIR = 'D:/Form5500_Downloads' if os.path.exists(DOWNLOAD_DIR): shutil.rmtree(DOWNLOAD_DIR) os.makedirs(DOWNLOAD_DIR, exist_ok=True) '''Function to download a single form using ACK ID''' def download_form(ackid): # Create a unique temp directory for each Chrome instance to avoid cache conflicts with tempfile.TemporaryDirectory() as temp_profile_dir: # Setting Chrome preferences chromeOptions = webdriver.ChromeOptions() prefs = {"download.default_directory": DOWNLOAD_DIR} chromeOptions.add_experimental_option("prefs", prefs) # Fixes for process cleanup and cache conflicts chromeOptions.add_argument(f"--user-data-dir={temp_profile_dir}") # Unique profile per instance chromeOptions.add_argument("--no-sandbox") # Fixes permission issues in multiprocessing chromeOptions.add_argument("--disable-dev-shm-usage") # Avoids shared memory limits chromeOptions.add_argument("--headless=new") # Optional: Run without visible windows (faster) chromeOptions.add_argument("--disable-gpu") # Disables GPU to avoid shader cache issues path_to_chromedriver = 'D:/401/401k/chromedriver_2.35.exe' browser = None try: browser = webdriver.Chrome(executable_path=path_to_chromedriver, options=chromeOptions) browser.implicitly_wait(10) # Longer implicit wait to handle slow page loads # Open ERISA website url = 'https://www.efast.dol.gov/portal/app/disseminate?execution=e1s4#' browser.get(url) # Search for a form using ACK ID browser.find_element(By.CSS_SELECTOR, '#ackId').send_keys(ackid) browser.find_element(By.CSS_SELECTOR, '.ui-icon-search').click() # Check if the form exists - if NOT, exit the function try: browser.find_element(By.CSS_SELECTOR, '#form\:filingTreeTable\:0\:einLnk').click() except: print(f"No form found for ACK ID: {ackid}") return # Exit early if no form exists # Wait until downloaded and rename using ackid print(f"Processing: {ackid}") download_path = os.path.join(DOWNLOAD_DIR, 'filing.pdf') timeout = 60 # 1 minute timeout to avoid hanging forever start_time = time.time() while not os.path.exists(download_path): if time.time() - start_time > timeout: print(f"Download timed out for ACK ID: {ackid}") return time.sleep(2) # Shorter sleep interval for responsiveness # Handle case where file is still being written (Chrome sometimes locks the file) time.sleep(1) os.rename(download_path, os.path.join(DOWNLOAD_DIR, f"{ackid}.pdf")) print(f"Successfully downloaded: {ackid}") except Exception as e: print(f"Error processing {ackid}: {str(e)}") finally: # Ensure browser is closed no matter what happens if browser: browser.quit() # Force kill any remaining Chrome processes (extra safety) os.system(f"taskkill /im chrome.exe /f /t") def main(): # Get list of ackids from csv file df = pd.read_csv('D:/401/401k/F_SCH_H_2015_latest.csv', usecols=[0], nrows=10000) return df['ACK_ID'].tolist() if __name__ == '__main__': ackid_list = main() # Use a smaller pool size if you run into resource issues (10 might be too many for your system) with Pool(5) as p: p.map(download_form, ackid_list)
Key Fixes Breakdown
- Variable Scope Fix: Moved
ackid_listretrieval tomain()and called it before starting the Pool, so the variable exists when needed. - Unique Chrome Profiles: Used
tempfile.TemporaryDirectory()to create a unique user data directory for each Chrome instance—this eliminates the shader cache conflict entirely. - Robust Cleanup: Added a
try/finallyblock to ensurebrowser.quit()runs no matter what exceptions occur. Also added ataskkillcommand as a last resort to kill any stuck Chrome processes. - Timeout for Downloads: Added a 1-minute timeout to avoid infinite loops if a download gets stuck.
- Chrome Stability Flags: Added
--no-sandbox,--disable-dev-shm-usage, and--disable-gputo prevent crashes and process leaks in multiprocessing environments. - Optional Headless Mode: Enabled
--headless=newto run Chrome without visible windows—this uses less system resources and avoids cluttering your screen with dozens of windows. You can remove this line if you need to see the browser for debugging.
Additional Tips
- Adjust Pool Size: 10 concurrent Chrome instances might overwhelm your system's memory/CPU. Start with 5-6 and increase if your system can handle it.
- Update ChromeDriver: Your ChromeDriver version (2.35) is very old—make sure it matches your installed Chrome version. Outdated drivers can cause unexpected crashes and compatibility issues.
- Check Download Locking: Sometimes Chrome keeps the
filing.pdffile locked for a second after download—adding a shorttime.sleep(1)before renaming prevents "file in use" errors.
内容的提问来源于stack exchange,提问作者Ankit Srivastava

