You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Cloud Functions中用Playwright实现Python网页抓取遇阻

解决GCP中Playwright无Chromium可执行文件的问题及Cloud Run部署指南

问题根源

Cloud Functions的Python运行环境默认不会预装Playwright所需的Chromium浏览器二进制文件。部署时仅会安装requirements.txt中声明的Python依赖,但playwright install是独立的命令,用于下载浏览器组件,因此运行时会因找不到可执行文件抛出错误,导致500内部服务器异常。

解决方案:Docker部署到Cloud Run

Cloud Run支持自定义Docker镜像,可完整包含Playwright及浏览器依赖,是更稳定的方案,以下是具体操作步骤:

1. 编写Dockerfile

创建包含Python环境、系统依赖、Playwright及Chromium的镜像:

# 基于官方Python slim镜像,减少体积
FROM python:3.11-slim

WORKDIR /app

# 安装Playwright运行所需的系统库
RUN apt-get update && apt-get install -y --no-install-recommends \
    libnss3 libatk1.0-0 libatk-bridge2.0-0 libcups2 libdrm2 libxkbcommon0 \
    libxcomposite1 libxdamage1 libxfixes3 libxrandr2 libgbm1 libasound2 \
    && rm -rf /var/lib/apt/lists/*

# 复制依赖清单并安装Python包
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

# 安装Chromium浏览器及对应依赖
RUN playwright install chromium && playwright install-deps chromium

# 复制业务代码
COPY main.py .

# 设置Cloud Run默认端口
ENV PORT=8080

# 启动HTTP服务(使用FastAPI+Uvicorn)
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "${PORT}"]

2. 编写requirements.txt

声明所需Python依赖:

fastapi
uvicorn
playwright

3. 编写Python服务代码(main.py)

实现GET请求触发截图的功能:

from fastapi import FastAPI, Response
from playwright.sync_api import sync_playwright
import logging

app = FastAPI()
logging.basicConfig(level=logging.INFO)

@app.get("/")
def capture_google_screenshot():
    try:
        with sync_playwright() as p:
            # 启动无头浏览器,添加适配Cloud Run的参数
            browser = p.chromium.launch(
                headless=True,
                args=["--no-sandbox", "--disable-dev-shm-usage"]
            )
            page = browser.new_page()
            page.goto("https://www.google.com", wait_until="networkidle")
            screenshot = page.screenshot(full_page=True)
            browser.close()
            return Response(content=screenshot, media_type="image/png")
    except Exception as e:
        logging.error(f"截图失败: {str(e)}")
        return {"error": "截图生成失败"}, 500

4. 构建并推送镜像到GCR

替换命令中的PROJECT_ID为你的GCP项目ID:

# 配置Docker访问GCR权限
gcloud auth configure-docker

# 构建镜像
docker build -t gcr.io/PROJECT_ID/playwright-screenshot:v1 .

# 推送镜像到Google容器注册表
docker push gcr.io/PROJECT_ID/playwright-screenshot:v1

5. 部署到Cloud Run

命令行方式:

gcloud run deploy playwright-screenshot \
    --image gcr.io/PROJECT_ID/playwright-screenshot:v1 \
    --platform managed \
    --region us-central1 \
    --allow-unauthenticated \
    --cpu 1 \
    --memory 512Mi

控制台方式:

  1. 打开GCP控制台的Cloud Run页面,点击「创建服务」
  2. 在「容器镜像URL」中选择刚才推送的镜像
  3. 设置端口为8080,勾选「允许未认证调用」(若需公开访问)
  4. 配置CPU为1vCPU、内存512MB,点击「创建」

最佳实践

  • 浏览器启动优化:添加--no-sandbox和--disable-dev-shm-usage参数,适配Cloud Run的沙箱环境与内存限制
  • 资源配置:浏览器运行需要足够资源,建议至少配置1vCPU和512MB内存,避免启动失败
  • 镜像体积优化:使用slim镜像、清理apt缓存、使用--no-cache-dir安装Python依赖,减少镜像大小
  • 错误处理:添加异常捕获与日志记录,便于排查问题,同时返回友好的HTTP状态码
  • 日志排查:通过Cloud Logging查看运行日志,定位浏览器启动、页面加载等环节的异常

内容的提问来源于stack exchange,提问作者Wesley Jeftha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 17:33:22