在AWS Lambda的Selenium Docker容器中读取Marionette端口失败
解决AWS Lambda容器中Firefox Marionette端口读取失败问题
问题背景
使用以下Dockerfile和app.py构建AWS Lambda容器镜像,运行时抛出TimeoutException: Message: Failed to read marionette port错误,此前尝试Chrome也遇到安装和交互问题。
原Dockerfile
FROM public.ecr.aws/lambda/python:3.9 RUN yum update -y RUN yum install -y \ Xvfb \ wget \ gtk3 \ dbus-glib \ libpci \ unzip \ gcc \ openssl-devel \ zlib-devel \ libffi-devel \ libgtk-3-0 \ alsa-lib-devel \ # I think needed for pandas libxml2 \ libxml2-devel \ g++ \ yum -y clean all RUN yum -y groupinstall development WORKDIR /opt RUN wget -O- "https://download.mozilla.org/?product=firefox-latest-ssl&os=linux64&lang=en-US" | tar -jx -C /usr/local/ # Borrowed from here: https://github.com/aws-samples/container-web-scraper-example/blob/master/code/Dockerfile RUN ln -s /usr/local/firefox/firefox /usr/bin/firefox RUN wget https://github.com/mozilla/geckodriver/releases/download/v0.31.0/geckodriver-v0.31.0-linux64.tar.gz RUN tar -xf geckodriver-v0.31.0-linux64.tar.gz RUN ls -lta RUN rm geckodriver-v0.31.0-linux64.tar.gz RUN chmod +x geckodriver RUN export DISPLAY=:99 RUN Xvfb -ac -nolisten inet6 :99 & WORKDIR /var/task # Install selenium COPY lambda_reqs.txt . RUN pip3 install -r lambda_reqs.txt # Copy lambda's main script COPY app.py . CMD ["app.lambda_handler"]
原app.py
import os import boto3 from io import StringIO from selenium import webdriver from selenium.webdriver.firefox.service import Service from selenium.webdriver import ActionChains from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.webdriver.firefox.options import Options import pandas as pd executable_path = '/opt/geckodriver' options = Options() options.headless = True options.add_argument("--no-sandbox") options.add_argument("--single-process") options.add_argument("--disable-dev-shm-usage") driver = webdriver.Firefox(options=options, executable_path='/opt/geckodriver' service_log_path=os.path.devnull, ) def lambda_handler(event, context): """ Invoke AWS Lambda Function :param event: :param context: :return: """ # More sample code than actual driver.get("https://www.google.com/") element_text = driver.page_source
错误信息
[ERROR] TimeoutException: Message: Failed to read marionette port Traceback (most recent call last): File "/var/lang/lib/python3.9/importlib/__init__.py", line 127, in import_module return _bootstrap._gcd_import(name[level:], package, level) File "<frozen importlib._bootstrap>", line 1030, in _gcd_import File "<frozen importlib._bootstrap>", line 1007, in _find_and_load File "<frozen importlib._bootstrap>", line 986, in _find_and_load_unlocked File "<frozen importlib._bootstrap>", line 680, in _load_unlocked File "<frozen importlib._bootstrap_external>", line 850, in exec_module File "<frozen importlib._bootstrap>", line 228, in _call_with_frames_removed File "/var/task/app.py", line 19, in <module> driver = webdriver.Firefox(options=options, executable_path='/opt/geckodriver', service_log_path=os.path.devnull) File "/var/lang/lib/python3.9/site-packages/selenium/webdriver/firefox/webdriver.py", line 177, in __init__ RemoteWebDriver.__init__( File "/var/lang/lib/python3.9/site-packages/selenium/webdriver/remote/webdriver.py", line 275, in __init__ self.start_session(capabilities, browser_profile) File "/var/lang/lib/python3.9/site-packages/selenium/webdriver/remote/webdriver.py", line 365, in start_session response = self.execute(Command.NEW_SESSION, parameters) File "/var/lang/lib/python3.9/site-packages/selenium/webdriver/remote/webdriver.py", line 430, in execute self.error_handler.check_response(response) File "/var/lang/lib/python3.9/site-packages/selenium/webdriver/remote/errorhandler.py", line 247, in check_response raise exception_class(message, screen, stacktrace)
解决方案
1. 修复Dockerfile中的环境变量和依赖问题
- 用
ENV DISPLAY=:99替代RUN export DISPLAY=:99,确保环境变量保留到容器运行时 - 构建时启动的Xvfb进程不会在容器运行时存活,移除构建时的Xvfb启动命令,改为在代码中启动
- 锁定Firefox版本为稳定ESR版,确保与geckodriver v0.31.0兼容
- 清理yum缓存,减小镜像体积
修改后的Dockerfile:
FROM public.ecr.aws/lambda/python:3.9 # 设置持久化环境变量 ENV DISPLAY=:99 RUN yum update -y && \ yum install -y \ Xvfb \ wget \ gtk3 \ dbus-glib \ libpci \ unzip \ gcc \ openssl-devel \ zlib-devel \ libffi-devel \ libgtk-3-0 \ alsa-lib-devel \ libxml2 \ libxml2-devel \ g++ && \ yum -y groupinstall development && \ yum -y clean all && \ rm -rf /var/cache/yum WORKDIR /opt # 下载Firefox 102ESR,与geckodriver v0.31.0稳定兼容 RUN wget -O- "https://download.mozilla.org/?product=firefox-102.15.0esr-ssl&os=linux64&lang=en-US" | tar -jx -C /usr/local/ RUN ln -s /usr/local/firefox/firefox /usr/bin/firefox # 安装指定版本geckodriver RUN wget https://github.com/mozilla/geckodriver/releases/download/v0.31.0/geckodriver-v0.31.0-linux64.tar.gz && \ tar -xf geckodriver-v0.31.0-linux64.tar.gz && \ rm geckodriver-v0.31.0-linux64.tar.gz && \ chmod +x geckodriver WORKDIR /var/task COPY lambda_reqs.txt . RUN pip3 install -r lambda_reqs.txt --no-cache-dir COPY app.py . CMD ["app.lambda_handler"]
2. 调整app.py的驱动初始化逻辑
- 不在模块级别提前初始化driver,避免Lambda冷启动时Xvfb未就绪
- 在
lambda_handler内部启动Xvfb,等待就绪后再初始化Firefox - 显式指定Marionette端口,避免自动分配冲突
- 执行完成后关闭driver和Xvfb进程,释放资源
修改后的app.py:
import os import subprocess import time import boto3 from io import StringIO from selenium import webdriver from selenium.webdriver.firefox.service import Service from selenium.webdriver import ActionChains from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By from selenium.webdriver.firefox.options import Options import pandas as pd def lambda_handler(event, context): # 启动Xvfb虚拟显示服务 xvfb_process = subprocess.Popen(['Xvfb', ':99', '-ac', '-nolisten', 'inet6']) time.sleep(1) # 给服务启动预留时间 options = Options() options.headless = True options.add_argument("--no-sandbox") options.add_argument("--single-process") options.add_argument("--disable-dev-shm-usage") # 显式指定Marionette端口,解决端口读取超时问题 options.add_argument('--marionette-port=2828') driver = None try: driver = webdriver.Firefox( options=options, executable_path='/opt/geckodriver', service_log_path=os.path.devnull ) driver.get("https://www.google.com/") element_text = driver.page_source print("页面加载成功") return {"statusCode": 200, "body": "页面抓取成功"} finally: # 确保资源被正确释放 if driver: driver.quit() xvfb_process.terminate()
3. 锁定依赖版本兼容性
在lambda_reqs.txt中指定与环境匹配的依赖版本:
selenium==4.4.3 pandas==1.5.3 boto3==1.28.0
Chrome替代方案(解决安装与交互问题)
若仍想使用Chrome,可按以下方式调整:
- 安装Chrome稳定版而非旧版本,适配Amazon Linux 2环境
- 使用Chrome 112+的新无头模式
--headless=new,避免交互异常 - 自动匹配chromedriver与Chrome版本
Dockerfile中Chrome安装片段:
# 安装Chrome稳定版 RUN wget https://dl.google.com/linux/direct/google-chrome-stable_current_x86_64.rpm && \ yum install -y google-chrome-stable_current_x86_64.rpm && \ rm google-chrome-stable_current_x86_64.rpm # 下载对应版本的chromedriver RUN CHROME_VERSION=$(google-chrome --version | grep -oP '\d+\.\d+\.\d+\.\d+') && \ wget https://storage.googleapis.com/chrome-for-testing-public/$CHROME_VERSION/linux64/chromedriver-linux64.zip && \ unzip chromedriver-linux64.zip && \ mv chromedriver-linux64/chromedriver /opt/chromedriver && \ chmod +x /opt/chromedriver && \ rm -rf chromedriver-linux64.zip chromedriver-linux64
内容的提问来源于stack exchange,提问作者jClean
相关产品推荐
相关产品推荐

