You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

AWS Lambda部署Python Selenium爬虫遇只读文件系统错误求助

解决方案

1. 修改Dockerfile:预安装ChromeDriver并精简镜像

在镜像构建阶段就完成ChromeDriver的下载与配置,避免运行时在Lambda只读目录下执行安装操作,同时优化镜像体积:

FROM python:3.10-slim
WORKDIR /app

COPY requirements.txt .

# 安装依赖、Chrome及对应版本的ChromeDriver
RUN apt-get update && apt-get install -y --no-install-recommends wget unzip google-chrome-stable && \
    # 获取当前Chrome版本,匹配对应ChromeDriver
    CHROME_VERSION=$(google-chrome --version | grep -Eo "[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+") && \
    wget https://storage.googleapis.com/chrome-for-testing-public/$CHROME_VERSION/linux64/chromedriver-linux64.zip && \
    unzip chromedriver-linux64.zip && \
    mv chromedriver-linux64/chromedriver /app/chromedriver && \
    chmod +x /app/chromedriver && \
    # 清理冗余文件缩小镜像
    rm -rf chromedriver-linux64.zip chromedriver-linux64 /var/lib/apt/lists/* && \
    apt-get clean && \
    pip install --no-cache-dir -r requirements.txt 

COPY . /app
CMD ["python", "LME_2.py"]

2. 修改driver.py:适配Lambda只读文件系统

指定ChromeDriver固定路径,同时将Chrome的所有可写操作指向Lambda唯一可写的/tmp目录:

import selenium 
from selenium import webdriver 
from selenium.webdriver.chrome.service import Service as ChromeService
import os

def return_driver(headless = True, logging = True):
    options = webdriver.ChromeOptions()
    options.add_argument('--no-sandbox')
    options.add_argument('--disable-dev-shm-usage')
    options.add_argument("--disable-gpu")    
    if headless:
        options.add_argument("--headless=new")  # 新版无头模式兼容性更好
    options.add_argument("--window-size=1920,1080")
    options.add_argument("--log-level=2")
    # 指定Chrome用户数据目录到/tmp,避免写入只读目录
    options.add_argument(f"--user-data-dir=/tmp/chrome-user-data")
    # 禁用自动更新及后台网络请求,避免权限问题
    options.add_argument("--disable-background-networking")
    options.add_argument("--disable-features=VizDisplayCompositor")
    
    try:
        # 使用预安装的ChromeDriver路径
        service = ChromeService(executable_path='/app/chromedriver')
        driver = webdriver.Chrome(service=service, options=options)
        print("Chromedriver started successfully.")
        return driver
    except Exception as e: 
        print(f"Error creating Chrome driver: {e}")
        raise

3. 可选:保留ChromeDriverManager的适配方案

如果需要自动适配Chrome版本,可配置ChromeDriverManager将缓存目录指向/tmp:

from webdriver_manager.chrome import ChromeDriverManager
from webdriver_manager.core.driver_cache import DriverCacheManager

# 替换原ChromeDriverManager调用部分
cache_manager = DriverCacheManager(cache_dir='/tmp/.wdm')
driver_path = ChromeDriverManager(cache_manager=cache_manager).install()
service = ChromeService(executable_path=driver_path)

额外注意事项

  • Lambda的/tmp目录默认最大存储空间为512MB,需确保爬虫运行时产生的临时文件不超过此限制;
  • 若镜像体积过大,可进一步使用python:3.10-alpine基础镜像(需额外安装Chrome依赖库)。

内容的提问来源于stack exchange,提问作者KAPIL BADOKAR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.02 15:15:27