远程机器运行Python爬虫时Chrome异常退出的问题求助
Hey there, let's work through this frustrating crash issue you're facing with your web/PDF scraping crawler on the remote machine. That "chrome exited abnormally" error is super common with headless Chrome setups on servers, and there are several tried-and-true fixes to get your crawler running reliably:
1. Match Chrome and ChromeDriver Versions Exactly
Selenium is picky about version compatibility—if your Chrome browser and ChromeDriver don't match, crashes are almost guaranteed.
First, check the versions on your remote machine:
# Check Chrome version google-chrome --version # Check ChromeDriver version chromedriver --version
If they don't line up (e.g., Chrome 114.x needs ChromeDriver 114.x), download the matching ChromeDriver, replace the existing one, and make sure it's executable:
chmod +x /path/to/your/chromedriver
2. Add Critical Headless Mode Arguments
Remote servers don't have a GUI, so you must run Chrome in headless mode—and adding these flags will prevent most unexpected crashes:
from selenium import webdriver from selenium.webdriver.chrome.options import Options def get_chrome_driver(): chrome_options = Options() # Enable updated headless mode for newer Chrome versions chrome_options.add_argument('--headless=new') # Disable sandbox (required on most Linux servers) chrome_options.add_argument('--no-sandbox') # Fix /dev/shm memory limitations chrome_options.add_argument('--disable-dev-shm-usage') # Disable GPU rendering (unnecessary in headless mode) chrome_options.add_argument('--disable-gpu') # Optional: Limit memory usage to prevent resource exhaustion chrome_options.add_argument('--memory-pressure-off') chrome_options.add_argument('--disable-features=VizDisplayCompositor') return webdriver.Chrome(options=chrome_options)
Here's why these matter:
--no-sandbox: Linux servers often restrict sandboxed processes, which Chrome relies on by default. Disabling it removes this barrier.--disable-dev-shm-usage: Many remote servers have a tiny/dev/shmpartition, which Chrome uses for temporary storage. This flag forces Chrome to use disk storage instead.
3. Monitor System Resources and Permissions
Crashes often happen when the remote machine runs out of memory or the crawler lacks proper permissions:
- Check memory usage: Run
htoportopto see if Chrome is eating up all available RAM. If so, optimize your crawler to quit the driver after each scrape (instead of keeping it open) or upgrade the server's memory. - Fix permissions: Ensure the user running the crawler has access to Chrome and ChromeDriver. Avoid running as root if possible, but if you must, keep the
--no-sandboxflag enabled. You can also adjust file permissions:chmod +x /usr/bin/google-chrome chmod +x /path/to/chromedriver
4. Add Error Handling to Auto-Restart the Driver
Since you're not monitoring the crawler full-time, adding exception handling to restart the driver when it crashes will keep your scraper running:
from selenium.common.exceptions import WebDriverException import time driver = get_chrome_driver() try: # Your scraping logic here driver.get("https://example.com") # ... rest of your code except WebDriverException as e: print(f"Driver crashed with error: {e}") # Clean up the crashed driver try: driver.quit() except: pass # Wait a few seconds before restarting time.sleep(5) driver = get_chrome_driver() # Resume your scraping logic
5. Upgrade Python and Selenium (Strongly Recommended)
You're using Python 2.7, which has been end-of-life since 2020. Selenium's support for Python 2.7 is outdated, and many newer Chrome features won't work properly. Upgrading to Python 3.8+ and installing the latest Selenium will eliminate a ton of compatibility issues:
# Install latest Selenium for Python 3 pip3 install --upgrade selenium
Bonus: Debug with Chrome Logs
If you still can't pinpoint the issue, enable Chrome logging to get detailed error messages:
chrome_options.add_argument('--enable-logging') chrome_options.add_argument('--v=1') # Adjust verbosity level (1-5) driver = webdriver.Chrome( options=chrome_options, service_log_path='/tmp/chromedriver_debug.log' )
Then check the log file:
cat /tmp/chromedriver_debug.log
This will show you exactly why Chrome exited—whether it's a missing library, resource limit, or configuration issue.
内容的提问来源于stack exchange,提问作者Neue1987

