如何在服务器上同时多次运行含阻塞代码的Python脚本(含Selenium场景)
Python服务器环境下并行运行阻塞Selenium任务的最优方案
针对你提到的「多次运行含阻塞代码的函数(尤其是Selenium启动Chrome的场景)」,结合服务器环境的特性,以下是几种实用方案及最优选择:
1. 多进程(multiprocessing)—— 小规模任务首选
Selenium的ChromeDriver受Python GIL(全局解释器锁)限制,多线程无法实现真正并行,而多进程能完全隔离每个Chrome实例,避免资源冲突,是服务器环境下最直接的方案。
示例代码:
from multiprocessing import Pool from selenium import webdriver from selenium.webdriver.chrome.options import Options def run_selenium_task(task_id): # 服务器环境必须配置Chrome无头模式及专属参数 chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") # 绕过系统安全沙箱 chrome_options.add_argument("--disable-dev-shm-usage") # 解决临时空间不足问题 driver = webdriver.Chrome(options=chrome_options) # 这里替换为你的阻塞业务逻辑 driver.get("https://example.com") print(f"任务 {task_id} 完成: {driver.title}") driver.quit() # 必须关闭实例,避免残留僵尸进程 if __name__ == "__main__": task_count = 5 # 并行任务数,建议不超过服务器CPU核心数的70% with Pool(processes=task_count) as pool: pool.map(run_selenium_task, range(task_count))
- 优势:代码简单、无额外依赖,完全实现并行,服务器兼容性强
- 注意:进程数不要超过服务器承载上限,否则会导致资源耗尽
2. 异步+进程池(asyncio + ProcessPoolExecutor)—— 混合IO场景适配
如果你的业务本身涉及异步操作(比如同时处理HTTP请求),可以用这种方式兼顾异步灵活性和并行能力。
示例代码:
import asyncio from concurrent.futures import ProcessPoolExecutor from selenium import webdriver from selenium.webdriver.chrome.options import Options def selenium_task(task_id): chrome_options = Options() chrome_options.add_argument("--headless=new") chrome_options.add_argument("--no-sandbox") chrome_options.add_argument("--disable-dev-shm-usage") driver = webdriver.Chrome(options=chrome_options) driver.get("https://example.com") result = f"任务 {task_id} 结果: {driver.title}" driver.quit() return result async def main(): task_count = 5 with ProcessPoolExecutor(max_workers=task_count) as executor: loop = asyncio.get_event_loop() tasks = [loop.run_in_executor(executor, selenium_task, i) for i in range(task_count)] results = await asyncio.gather(*tasks) for res in results: print(res) if __name__ == "__main__": asyncio.run(main())
- 优势:适配复杂混合IO场景,既保留异步的高效,又规避GIL实现并行
3. 容器化部署(Docker)—— 大规模长期任务最优解
如果需要长期稳定运行大量Selenium任务,容器化是最可靠的方案,能彻底隔离环境,避免版本依赖、进程泄漏等问题,还能弹性扩容。
简化示例(Dockerfile):
FROM python:3.11-slim # 安装Chrome及驱动 RUN apt-get update && apt-get install -y chromium chromium-driver WORKDIR /app COPY requirements.txt . RUN pip install --no-cache-dir selenium COPY your_selenium_script.py . CMD ["python", "your_selenium_script.py"]
- 优势:环境完全隔离,可通过Docker Compose/K8s批量管理任务,服务器压力大时直接扩容容器数量
- 适用场景:长期运行的爬虫、自动化测试等大规模任务
方案优先级总结
- 临时/小规模任务:选多进程,快速实现无额外成本
- 混合异步业务:选asyncio+进程池,兼顾灵活性与并行
- 长期/大规模任务:选Docker容器化,稳定性和可维护性最高
核心注意事项
- 服务器上必须保证Chrome与ChromeDriver版本严格匹配,否则会启动失败
- 所有Selenium任务必须显式调用
driver.quit(),避免残留大量Chrome进程耗尽资源 - 绝对不要用多线程(
threading)处理Selenium任务,GIL会导致无法真正并行,还可能引发实例冲突崩溃
内容的提问来源于stack exchange,提问作者Justin Battaglia
相关产品推荐
相关产品推荐

