Python multiprocessing执行复杂Selenium函数并行运行问题求助
Hey there! Let's work through getting your download and upload functions running in parallel—this is a common scenario with Selenium, and I’ve got a solid approach to share with you.
Key Pre-Requisite: Isolate WebDriver Instances
First off, a critical rule: Selenium's WebDriver is not thread-safe. You can’t share a single driver between your two functions when running them in parallel. Each function needs its own dedicated WebDriver instance to avoid race conditions and unexpected crashes.
Step-by-Step Implementation with threading
Since your functions are mostly IO-bound (waiting for page loads, file transfers), Python’s threading module is the right tool here—it’s lighter weight than multiprocessing and works perfectly for these kinds of tasks.
Here’s a concrete example tailored to your use case:
import threading import time from selenium import webdriver from selenium.webdriver.chrome.options import Options def run_download_workflow(): # Initialize a unique driver for this thread chrome_options = Options() # Configure download preferences (adjust path as needed) chrome_options.add_experimental_option( "prefs", {"download.default_directory": "/your/local/download/path"} ) # Add other options like headless mode if you don't need a GUI # chrome_options.add_argument("--headless=new") driver = webdriver.Chrome(options=chrome_options) try: # Your existing download logic with while loops and conditionals goes here while True: # Example condition: check if there are pending files to download if check_for_pending_downloads(): driver.get("https://your-download-site.com") # Perform download actions (click buttons, wait for files, etc.) execute_download_steps(driver) # Add a small delay to avoid spamming the server time.sleep(10) except Exception as e: print(f"Download thread encountered an error: {str(e)}") finally: # Always clean up the driver to free resources driver.quit() def run_upload_workflow(): # Initialize another unique driver for this thread chrome_options = Options() driver = webdriver.Chrome(options=chrome_options) try: # Your existing upload logic with while loops and conditionals goes here while True: # Example condition: check if new downloaded files exist if check_for_new_downloaded_files(): driver.get("https://your-upload-site.com") # Perform upload actions (select files, submit forms, etc.) execute_upload_steps(driver) time.sleep(10) except Exception as e: print(f"Upload thread encountered an error: {str(e)}") finally: driver.quit() # Helper functions (replace these with your actual logic) def check_for_pending_downloads(): # Logic to check if there are files to download return True def execute_download_steps(driver): # Your download steps here pass def check_for_new_downloaded_files(): # Logic to check for newly downloaded files return True def execute_upload_steps(driver): # Your upload steps here pass if __name__ == "__main__": # Create threads for each workflow download_thread = threading.Thread(target=run_download_workflow) upload_thread = threading.Thread(target=run_upload_workflow) # Start both threads download_thread.start() upload_thread.start() # Keep the main thread alive until both workflows finish (optional) download_thread.join() upload_thread.join()
Important Notes to Avoid Issues
- Cleanup with
finally: Always wrap your driver usage in atry/finallyblock to ensure the driver is quit properly, even if an error occurs. This prevents leftover browser processes from hogging memory. - Shared Data? Use Locks: If your download and upload functions need to share data (like a list of downloaded files), use a
threading.Lockto prevent race conditions. For example:# Shared variable and lock downloaded_files = [] file_lock = threading.Lock() # In download function, when a file is ready: with file_lock: downloaded_files.append("/path/to/new/file") # In upload function, when checking for files: with file_lock: files_to_upload = downloaded_files.copy() downloaded_files.clear() - Loop Exit Conditions: Make sure your
whileloops have a way to exit (e.g., a flag that gets set when you want to stop the workflow). Otherwise, the threads will run indefinitely.
Alternative: Multiprocessing
If you run into issues with threading (uncommon for Selenium IO-bound tasks), you can use Python’s multiprocessing module instead. The approach is similar—each process gets its own WebDriver instance—but processes are heavier weight. Threading is almost always the better choice here.
内容的提问来源于stack exchange,提问作者Jared Forth

