爬取万级网站时Selenium WebDriver随机无报错卡顿问题的排查与解决咨询
get() Hangs Without Errors Hey, I’ve dealt with exactly these kinds of random, silent hangs in Selenium before—super frustrating when you can’t reproduce them consistently! Let’s walk through the most likely causes and fixes to get your 10k-site crawl back on track.
Possible Causes & Fixes
1. Page Load Strategy Waiting for Unnecessary Resources
Your current setup uses Chrome’s default page load strategy, which waits for every resource (images, ads, third-party tracking scripts) to fully load. If a site has a broken ad script or slow external resource, Chrome might hang indefinitely—even with your 15-second timeout (sometimes the timeout fails to trigger for stuck network requests).
Fix: Switch to the eager load strategy, which stops waiting once the DOM is ready (no need to wait for non-critical assets):
chrome_options = webdriver.ChromeOptions() chrome_options.page_load_strategy = 'eager' # Add this line prefs = {"profile.default_content_setting_values.notifications": 2} chrome_options.add_experimental_option("prefs", prefs) driver_pcr = webdriver.Chrome(chrome_options=chrome_options) driver_pcr.set_page_load_timeout(15)
For even more control, you can use none (loads the page without waiting for any resources), but eager is usually the sweet spot for crawlers.
2. Mismatched ChromeDriver & Chrome Versions
This is a sneaky culprit—even a minor version gap between ChromeDriver and your installed Chrome can cause random compatibility issues like silent hangs. For example, Chrome 118 needs ChromeDriver 118.x, not 117.x.
Fix:
- Check your Chrome version (Settings > About Chrome)
- Download the exact matching ChromeDriver executable to align versions perfectly
- Replace your current ChromeDriver with the new one
3. Accumulated Browser Resource Bloat
Crawling 10k sites means your Chrome instance will pile up memory, cookies, and cached data over time. Eventually, this bloat causes slowdowns or random hangs as the browser runs out of resources.
Fix: Implement a driver restart cycle. For example, restart the browser every 50-100 sites to clear resources:
site_count = 0 max_sites_per_driver = 50 for url in url_list: site_count += 1 if site_count % max_sites_per_driver == 0: driver_pcr.quit() # Reinitialize the driver fresh chrome_options = webdriver.ChromeOptions() chrome_options.page_load_strategy = 'eager' prefs = {"profile.default_content_setting_values.notifications": 2} chrome_options.add_experimental_option("prefs", prefs) driver_pcr = webdriver.Chrome(chrome_options=chrome_options) driver_pcr.set_page_load_timeout(15) # Proceed with your get() and scraping logic
You can also add --headless=new to your Chrome options to cut down on RAM/CPU usage (headless mode uses far fewer resources than a visible browser window).
4. Network Fluctuations or Stuck Connections
Temporary network blips can cause get() requests to hang without triggering the timeout, even if the site loads fine on retry. Adding retry logic mitigates this.
Fix: Wrap your get() call in a retry loop with timeout handling:
from selenium.common.exceptions import TimeoutException max_retries = 3 for attempt in range(max_retries): try: driver_pcr.get(f"https://{url}") break # Success, exit retry loop except TimeoutException: if attempt == max_retries - 1: print(f"Failed to load {url} after {max_retries} attempts") continue # Skip to next site # Refresh the driver for the next attempt driver_pcr.refresh()
5. Problematic Chrome Features
Experimental Chrome features like GPU acceleration or sandboxing can cause random hangs, especially on specific OS configurations.
Fix: Disable these features via Chrome options:
chrome_options.add_argument('--disable-gpu') chrome_options.add_argument('--no-sandbox') # Helps on Linux, but can fix Windows issues too chrome_options.add_argument('--disable-dev-shm-usage') # Reduces memory pressure on Linux
6. Lingering ChromeDriver Processes
If your script crashes or exits unexpectedly, ChromeDriver processes might linger in the background, eating up system resources and causing future hangs.
Fix: Always call driver_pcr.quit() (not just close()) when done, and use try/finally to ensure cleanup:
try: # Your crawl logic here finally: driver_pcr.quit()
On Linux/macOS, you can manually kill leftover processes with pkill chromedriver; on Windows, use Task Manager to end Chrome/ChromeDriver processes if needed.
Final Tips
- Log every site you crawl and any hangs—this can help spot patterns (e.g., hangs only on certain TLDs or sites with specific scripts)
- Test fixes with a small subset of URLs first before scaling to 10k sites
内容的提问来源于stack exchange,提问作者Alain

