Databricks环境下ChromeDriver与Selenium配置问题求助
问题描述
我编写了一个运行耗时较长的复杂网页爬虫脚本,希望部署到Databricks集群以摆脱本地服务器依赖,但环境配置始终失败,尝试多个方案均未解决。
当前配置代码:
%pip install selenium %pip install chromedriver %pip install webdriver_manager %pip install beautifulsoup4 import pandas as pd from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options from selenium.webdriver.support.ui import Select from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.common.action_chains import ActionChains from selenium.webdriver.support import expected_conditions as EC from webdriver_manager.chrome import ChromeDriverManager from webdriver_manager.core.utils import ChromeType import time from bs4 import BeautifulSoup import pickle as pkl service=Service(ChromeDriverManager().install()) driver = webdriver.Chrome(service=service)
尝试过程中遇到两类错误:
"databricks" "WebDriverException: Message: unknown error: cannot find Chrome binary"
WebDriverException: Message: unknown error: Chrome failed to start: exited abnormally.
(unknown error: DevToolsActivePort file doesn't exist)
(The process started from chrome location /usr/bin/chromium-browser is no longer running, so ChromeDriver is assuming that Chrome has crashed.)
Stacktrace:
#0 0x55c30eb24d93......
目前卡在第二类错误,已尝试多个相关方案但无效。
核心问题分析
Databricks集群默认未安装Chrome浏览器,且运行环境为无界面模式,直接启动Chrome会因缺少图形环境、端口配置冲突、沙箱机制限制等问题导致崩溃。
完整配置步骤
1. 安装Chrome浏览器及依赖
在Databricks Notebook中执行以下Shell命令(针对Ubuntu系统集群):
%sh sudo apt-get update sudo apt-get install -y chromium-browser chromium-chromedriver # 建立软链接,确保ChromeDriver路径可被Selenium识别 sudo ln -s /usr/lib/chromium-browser/chromedriver /usr/bin/chromedriver
2. 配置Selenium关键参数
必须添加无界面运行、禁用沙箱等参数,解决DevToolsActivePort和异常退出问题:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.chrome.options import Options # 配置Chrome运行选项 chrome_options = Options() chrome_options.add_argument("--headless=new") # 新版Chrome推荐的无界面模式 chrome_options.add_argument("--no-sandbox") # 禁用沙箱,Databricks环境必需 chrome_options.add_argument("--disable-dev-shm-usage") # 解决内存不足导致的崩溃 chrome_options.add_argument("--remote-debugging-port=9222") # 指定调试端口,避免端口冲突 chrome_options.binary_location = "/usr/bin/chromium-browser" # 指定Chrome二进制文件路径 # 初始化WebDriver service = Service("/usr/bin/chromedriver") driver = webdriver.Chrome(service=service, options=chrome_options)
3. 验证环境有效性
添加简单测试代码确认配置成功:
driver.get("https://www.example.com") print(driver.title) driver.quit()
4. 可选:用webdriver-manager自动匹配版本
若需自动管理ChromeDriver版本,需指定Chrome类型为CHROMIUM,配合上述选项使用:
from webdriver_manager.chrome import ChromeDriverManager from webdriver_manager.core.utils import ChromeType service = Service(ChromeDriverManager(chrome_type=ChromeType.CHROMIUM).install()) driver = webdriver.Chrome(service=service, options=chrome_options)
关键注意事项
- 若集群非Ubuntu系统,需调整浏览器安装命令
--no-sandbox和--disable-dev-shm-usage为Databricks环境下运行Chrome的必需参数,不可省略- 若出现版本不匹配问题,可通过
%sh chromium-browser --version查看浏览器版本,手动下载对应ChromeDriver
内容的提问来源于stack exchange,提问作者user11781

