You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

部署集成Selenium的Flask REST应用至生产环境的最佳方案咨询

集成Selenium的Flask REST API生产部署方案及并发最佳实践

一、解决现有部署问题

1. AWS Ubuntu + Gunicorn 下Selenium窗口尺寸问题

问题源于无头模式默认窗口过小,导致页面元素加载异常。启动Chrome时显式配置参数即可解决:

from selenium import webdriver
from selenium.webdriver.chrome.options import Options

def init_chrome_driver():
    chrome_options = Options()
    chrome_options.add_argument("--headless=new")  # 新版无头模式,更贴近正常浏览器行为
    chrome_options.add_argument("--no-sandbox")  # 适配Ubuntu环境权限限制
    chrome_options.add_argument("--disable-dev-shm-usage")  # 避免临时空间不足
    chrome_options.add_argument("--window-size=1920,1080")  # 设置标准屏幕尺寸
    chrome_options.add_argument("--disable-gpu")  # 禁用GPU加速,适配无头环境
    driver = webdriver.Chrome(options=chrome_options)
    return driver

同时需确保Ubuntu上安装的Chrome与ChromeDriver版本严格匹配。

2. AWS Windows Server + Uvicorn 的ASGI错误

Flask是WSGI应用,Uvicorn默认运行ASGI应用,直接部署会报错,两种修复方式:

  • 方式一:用asgiref适配为ASGI应用
from flask import Flask
from asgiref.wsgi import WSGIToASGI

app = Flask(__name__)
# 此处添加你的Flask路由逻辑

asgi_app = WSGIToASGI(app)

if __name__ == "__main__":
    import uvicorn
    uvicorn.run(asgiref_app, host="0.0.0.0", port=8000)
  • 方式二:改用Windows友好的WSGI服务器Waitress
pip install waitress

启动命令:

waitress-serve --host=0.0.0.0 --port=8000 your_app:app

二、生产环境推荐部署架构

1. 首选:Linux(Ubuntu/CentOS)环境

Linux稳定性与资源利用率更高,推荐架构:

  • 反向代理:Nginx(处理静态资源、负载均衡、SSL证书)
  • WSGI服务器:Gunicorn(管理Flask进程,建议工作进程数设为CPU核心数*2+1)
  • 进程监控:Supervisor(实时监控Gunicorn进程,异常自动重启)
  • Selenium依赖:安装Chrome稳定版+对应ChromeDriver,配置为无头模式

2. Windows Server环境(仅业务必须时使用)

若必须用Windows,推荐:

  • WSGI服务器:Waitress(替代Gunicorn,适配Windows环境)
  • 进程管理:注册为Windows系统服务,实现开机自启
  • Selenium配置:Chrome安装在默认路径,ChromeDriver添加到系统PATH,启用无头模式

三、多请求并发最佳实践

Selenium是单线程阻塞工具,直接在Flask请求线程中调用会导致请求排队,并发性能极差,需做隔离与异步处理:

1. 异步任务队列架构(核心方案)

将爬取任务从Flask请求线程剥离,放入异步队列,Flask仅负责接收请求、返回任务ID,客户端通过轮询获取结果:

  • 采用Celery + Redis/RabbitMQ作为任务队列
  • 示例配置:
# celery_config.py
from celery import Celery

celery = Celery(
    'crawl_tasks',
    broker='redis://localhost:6379/0',
    backend='redis://localhost:6379/0'
)

@celery.task
def execute_crawl(url):
    driver = init_chrome_driver()
    try:
        driver.get(url)
        # 此处添加爬取逻辑
        result = {"page_title": driver.title, "page_source": driver.page_source}
        return {"status": "success", "data": result}
    finally:
        driver.quit()

Flask路由示例:

from flask import jsonify, request
from celery_config import execute_crawl

@app.route('/api/crawl', methods=['POST'])
def submit_crawl():
    url = request.json.get('url')
    if not url:
        return jsonify({"error": "URL参数必填"}), 400
    task = execute_crawl.delay(url)
    return jsonify({"task_id": task.id}), 202

@app.route('/api/crawl/result/<task_id>')
def get_crawl_result(task_id):
    task = execute_crawl.AsyncResult(task_id)
    if task.state == 'PENDING':
        return jsonify({"status": "爬取中"})
    elif task.state == 'SUCCESS':
        return jsonify(task.result)
    else:
        return jsonify({"status": "失败", "error": str(task.info)}), 500

2. 资源池优化

若暂不使用任务队列,可提前初始化多个ChromeDriver实例,用连接池管理,避免每次请求启动新浏览器(启动开销极大):

  • 自定义连接池限制同时运行的Driver数量,防止服务器资源耗尽
  • 注意:每个Driver实例线程不安全,需确保同一时间仅一个线程使用

3. 并发控制

  • 限制Celery Worker数量(建议等于CPU核心数),避免Chrome进程过多导致资源耗尽
  • Nginx配置连接数限制,防止恶意请求压垮服务器
  • 对爬取接口添加频率限制,平衡目标网站反爬压力与自身服务器负载

内容的提问来源于stack exchange,提问作者Navitas28

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 07:32:06