Mac M2构建Python+Selenium Docker镜像部署Cloud Run遇Chrome缺失问题
解决Docker容器中Selenium找不到Google Chrome的问题
在Mac M2本地开发的Selenium爬虫运行正常,但构建Docker镜像部署到Google Cloud Run时,持续报错google-chrome: not found,多次调整Chrome和Selenium版本仍未解决。以下是针对性的解决方案:
1. 修改Dockerfile,安装Google Chrome及依赖
原Dockerfile未包含Chrome浏览器,这是核心问题。替换为以下Dockerfile,确保镜像中安装官方稳定版Chrome及运行所需依赖:
FROM python:3.8-slim # 安装Google Chrome及必要依赖 RUN apt-get update && apt-get install -y \ wget \ gnupg \ ca-certificates \ libnss3 \ libgconf-2-4 \ libxss1 \ fonts-liberation \ && wget -q -O - https://dl-ssl.google.com/linux/linux_signing_key.pub | apt-key add - \ && echo "deb [arch=amd64] http://dl.google.com/linux/chrome/deb/ stable main" >> /etc/apt/sources.list.d/google-chrome.list \ && apt-get update && apt-get install -y google-chrome-stable \ && apt-get clean \ && rm -rf /var/lib/apt/lists/* # 安装Python依赖 COPY requirements.txt . RUN pip install --no-cache-dir -r requirements.txt ENV APP_HOME /app WORKDIR $APP_HOME COPY . . CMD exec gunicorn --bind :$PORT --workers 1 --threads 8 main:app
- 使用
python:3.8-slim镜像减小体积,同时保证基础环境完整 - 通过官方源安装Chrome,避免版本不兼容问题
- 清理安装缓存,减少镜像大小
2. 调整requirements.txt
移除手动指定的chromedriver-binary,让webdriver-manager自动匹配Chrome版本,避免版本冲突:
Flask==1.0.2 gunicorn==19.9.0 selenium==4.11.2 google-cloud-bigquery==3.11.4 beautifulsoup4==4.10.0 webdriver-manager==4.0.0
3. 优化Python代码的Chrome配置
添加容器环境下运行Chrome的必要参数,避免无头模式报错:
def crawl_ag_grid(request): login_url = 'https://xxxxxx.com/login' target_url = 'https://xxxxx.com/dashboards/25' options = Options() # 容器环境无头模式必备配置 options.add_argument('--headless=new') # Selenium 4.8+推荐的稳定无头模式 options.add_argument('--no-sandbox') options.add_argument('--disable-dev-shm-usage') options.add_argument('--disable-gpu') options.add_argument('--window-size=1920,1080') # 自动匹配ChromeDriver版本 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()), options=options) try: driver.get(login_url) # 执行登录操作 email = driver.find_element(By.NAME,'email') password = driver.find_element(By.NAME,'password') email.send_keys('xxxxx@xx.com') password.send_keys('xxxxxxxxxx') login = driver.find_element(By.ID,'login-submit') login.submit() # 后续爬取逻辑... finally: driver.quit() # 确保关闭浏览器,释放容器资源
--no-sandbox和--disable-dev-shm-usage解决容器环境下Chrome的权限与内存限制问题- 添加
finally块保证浏览器资源被正确释放
完成以上修改后,重新构建Docker镜像并部署到Google Cloud Run即可正常运行爬虫。
内容的提问来源于stack exchange,提问作者bernard
相关产品推荐
相关产品推荐

