You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将Python Selenium脚本部署到服务器并通过API实现输入输出?

问题

我有一个Python Selenium脚本,它接收输入URL、爬取特定网站并生成输出。我希望将该脚本部署到服务器上,以便通过API发送输入,脚本将利用该输入爬取网站并通过API返回输出。我正在构建一个Web应用,用户可在此输入信息,脚本爬取数据后将结果返回给用户。

我曾在Heroku和多个VPS上运行该脚本,但尚未解决如何通过API调用脚本并返回数据的问题。我需要将其部署到服务器,以便随时通过发送输入来调用脚本,输入将启动爬虫、爬取数据并返回结果。以下是我想要部署的脚本通用形式:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

# 初始化浏览器驱动
def bot_setup(headless=False):
    options = webdriver.ChromeOptions()
    options.add_argument("--no-sandbox")
    options.add_argument("--disable-dev-shm-usage")
    options.add_argument("--disable-blink-features=AutomationControlled")
    options.add_experimental_option("useAutomationExtension", False)
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option("excludeSwitches", ["enable-logging"])
    options.add_argument("disable-infobars")
    if headless:
        options.add_argument("--headless")
    
    driver = webdriver.Chrome(
        service=Service(ChromeDriverManager().install()),
        options=options,
    )
    driver.implicitly_wait(10)
    return driver
    
def scrape(url):
    driver = bot_setup()
    driver.get(url)
    results = {}
    # 爬取逻辑写在这里
    return results # 需要通过API返回这些结果给前端或数据库

url = "https://example.com"
scrape(url) # 希望这个URL来自API请求

解决方案

1. 用Flask搭建API接口

用轻量Web框架Flask快速封装爬虫为API接口,修改后的完整代码如下:

from flask import Flask, request, jsonify
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

app = Flask(__name__)

def bot_setup(headless=True):  # 服务器运行必须启用无头模式
    options = webdriver.ChromeOptions()
    options.add_argument("--no-sandbox")
    options.add_argument("--disable-dev-shm-usage")
    options.add_argument("--disable-blink-features=AutomationControlled")
    options.add_experimental_option("useAutomationExtension", False)
    options.add_experimental_option("excludeSwitches", ["enable-automation"])
    options.add_experimental_option("excludeSwitches", ["enable-logging"])
    options.add_argument("disable-infobars")
    options.add_argument("--headless=new")  # 新版无头模式兼容性更好
    options.add_argument("--window-size=1920,1080")  # 避免元素定位异常

    driver = webdriver.Chrome(
        service=Service(ChromeDriverManager().install()),
        options=options,
    )
    driver.implicitly_wait(10)
    return driver

def scrape(url):
    driver = None
    try:
        driver = bot_setup()
        driver.get(url)
        results = {}
        # 替换为你的实际爬取逻辑,比如获取页面标题
        # results["page_title"] = driver.title
        return {"success": True, "data": results}
    except Exception as e:
        return {"success": False, "error": str(e)}
    finally:
        if driver:
            driver.quit()  # 强制关闭浏览器,避免服务器资源泄漏

# 定义API接口,接收POST请求
@app.route('/scrape', methods=['POST'])
def scrape_endpoint():
    # 解析请求中的JSON数据
    request_data = request.get_json()
    if not request_data or 'url' not in request_data:
        return jsonify({"success": False, "error": "请求缺少URL参数"}), 400
    
    target_url = request_data['url']
    scrape_result = scrape(target_url)
    return jsonify(scrape_result)

if __name__ == '__main__':
    app.run(host='0.0.0.0', port=5000)  # 允许外部访问API

2. 服务器环境配置

VPS(以Ubuntu为例)

  1. 安装Chrome浏览器:
sudo apt update
sudo apt install -y chromium-browser
  1. 安装依赖包:
pip install flask selenium webdriver-manager gunicorn
  1. 用gunicorn启动生产环境服务:
gunicorn --bind 0.0.0.0:5000 app:app

Heroku部署

  1. 添加Chrome相关Buildpack:
    在Heroku项目设置中添加两个Buildpack:
    • https://github.com/heroku/heroku-buildpack-chromedriver.git
    • https://github.com/heroku/heroku-buildpack-google-chrome.git
  2. 创建Procfile文件,内容:
web: gunicorn app:app
  1. 生成requirements.txt:
pip freeze > requirements.txt
  1. 提交代码到Heroku仓库完成部署。

3. 测试API

部署完成后,用curl或Postman测试接口:

curl -X POST -H "Content-Type: application/json" -d '{"url": "https://example.com"}' http://你的服务器IP:5000/scrape

4. 优化建议

  • 并发处理:如果存在多用户请求,可引入线程池或Celery实现异步爬取,避免接口阻塞。
  • 错误细化:针对不同异常(如网络超时、元素未找到)返回更具体的错误信息,便于前端处理。
  • 资源管控:定期清理无效的浏览器进程,避免服务器内存占用过高。
  • 结果缓存:对重复请求的URL缓存爬取结果,减少重复爬取消耗。

内容的提问来源于stack exchange,提问作者zayndotexe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 13:40:26