如何将Python Selenium脚本部署到服务器并通过API实现输入输出?
问题
我有一个Python Selenium脚本,它接收输入URL、爬取特定网站并生成输出。我希望将该脚本部署到服务器上,以便通过API发送输入,脚本将利用该输入爬取网站并通过API返回输出。我正在构建一个Web应用,用户可在此输入信息,脚本爬取数据后将结果返回给用户。
我曾在Heroku和多个VPS上运行该脚本,但尚未解决如何通过API调用脚本并返回数据的问题。我需要将其部署到服务器,以便随时通过发送输入来调用脚本,输入将启动爬虫、爬取数据并返回结果。以下是我想要部署的脚本通用形式:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager # 初始化浏览器驱动 def bot_setup(headless=False): options = webdriver.ChromeOptions() options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("useAutomationExtension", False) options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("excludeSwitches", ["enable-logging"]) options.add_argument("disable-infobars") if headless: options.add_argument("--headless") driver = webdriver.Chrome( service=Service(ChromeDriverManager().install()), options=options, ) driver.implicitly_wait(10) return driver def scrape(url): driver = bot_setup() driver.get(url) results = {} # 爬取逻辑写在这里 return results # 需要通过API返回这些结果给前端或数据库 url = "https://example.com" scrape(url) # 希望这个URL来自API请求
解决方案
1. 用Flask搭建API接口
用轻量Web框架Flask快速封装爬虫为API接口,修改后的完整代码如下:
from flask import Flask, request, jsonify from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager app = Flask(__name__) def bot_setup(headless=True): # 服务器运行必须启用无头模式 options = webdriver.ChromeOptions() options.add_argument("--no-sandbox") options.add_argument("--disable-dev-shm-usage") options.add_argument("--disable-blink-features=AutomationControlled") options.add_experimental_option("useAutomationExtension", False) options.add_experimental_option("excludeSwitches", ["enable-automation"]) options.add_experimental_option("excludeSwitches", ["enable-logging"]) options.add_argument("disable-infobars") options.add_argument("--headless=new") # 新版无头模式兼容性更好 options.add_argument("--window-size=1920,1080") # 避免元素定位异常 driver = webdriver.Chrome( service=Service(ChromeDriverManager().install()), options=options, ) driver.implicitly_wait(10) return driver def scrape(url): driver = None try: driver = bot_setup() driver.get(url) results = {} # 替换为你的实际爬取逻辑,比如获取页面标题 # results["page_title"] = driver.title return {"success": True, "data": results} except Exception as e: return {"success": False, "error": str(e)} finally: if driver: driver.quit() # 强制关闭浏览器,避免服务器资源泄漏 # 定义API接口,接收POST请求 @app.route('/scrape', methods=['POST']) def scrape_endpoint(): # 解析请求中的JSON数据 request_data = request.get_json() if not request_data or 'url' not in request_data: return jsonify({"success": False, "error": "请求缺少URL参数"}), 400 target_url = request_data['url'] scrape_result = scrape(target_url) return jsonify(scrape_result) if __name__ == '__main__': app.run(host='0.0.0.0', port=5000) # 允许外部访问API
2. 服务器环境配置
VPS(以Ubuntu为例)
- 安装Chrome浏览器:
sudo apt update sudo apt install -y chromium-browser
- 安装依赖包:
pip install flask selenium webdriver-manager gunicorn
- 用gunicorn启动生产环境服务:
gunicorn --bind 0.0.0.0:5000 app:app
Heroku部署
- 添加Chrome相关Buildpack:
在Heroku项目设置中添加两个Buildpack:https://github.com/heroku/heroku-buildpack-chromedriver.githttps://github.com/heroku/heroku-buildpack-google-chrome.git
- 创建
Procfile文件,内容:
web: gunicorn app:app
- 生成
requirements.txt:
pip freeze > requirements.txt
- 提交代码到Heroku仓库完成部署。
3. 测试API
部署完成后,用curl或Postman测试接口:
curl -X POST -H "Content-Type: application/json" -d '{"url": "https://example.com"}' http://你的服务器IP:5000/scrape
4. 优化建议
- 并发处理:如果存在多用户请求,可引入线程池或Celery实现异步爬取,避免接口阻塞。
- 错误细化:针对不同异常(如网络超时、元素未找到)返回更具体的错误信息,便于前端处理。
- 资源管控:定期清理无效的浏览器进程,避免服务器内存占用过高。
- 结果缓存:对重复请求的URL缓存爬取结果,减少重复爬取消耗。
内容的提问来源于stack exchange,提问作者zayndotexe
相关产品推荐
相关产品推荐

