You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何处理多线程Python脚本中卡住的Headless Chrome实例?

解决多线程Chrome超时卡顿问题

核心思路

给每个线程绑定对应的Chrome进程,同时在任务逻辑里加入超时判断——一旦任务运行超过设定时长,就强制终止关联的Chrome进程,再清理线程资源,避免卡顿扩散。

具体实现步骤

1. 安装依赖

打开终端运行以下命令,安装必要工具包:

pip install selenium psutil

2. 完整代码示例

代码注释已写清各部分作用,你只需替换自己的业务逻辑:

import threading
import time
import psutil
from selenium import webdriver
from selenium.webdriver.chrome.options import Options

# 配置Chrome启动参数,按需调整
def get_chrome_options():
    options = Options()
    # options.add_argument("--headless=new")  # 需要无头模式就取消注释
    options.add_argument("--no-sandbox")  # 解决权限问题
    options.add_argument("--disable-dev-shm-usage")  # 解决内存不足问题
    options.add_argument("--disable-gpu")  # 禁用GPU加速,减少卡顿
    return options

# 单个线程的任务逻辑
def browser_task(task_id, timeout_seconds):
    driver = None
    chrome_process = None
    start_time = time.time()  # 记录任务启动时间

    try:
        # 启动Chrome
        driver = webdriver.Chrome(options=get_chrome_options())
        # 获取Chrome主进程ID,方便后续查杀
        driver_pid = driver.service.process.pid
        chrome_process = psutil.Process(driver_pid)
        print(f"任务{task_id}启动,Chrome进程ID: {driver_pid}")

        # --------------------------
        # 这里替换成你的实际业务代码
        # 比如访问网页、点击元素、数据抓取等
        driver.get("https://example.com")
        # 模拟业务执行,你可以删掉这段换成自己的逻辑
        while time.time() - start_time < timeout_seconds:
            if driver.title == "Example Domain":
                print(f"任务{task_id}执行完成")
                break
            time.sleep(1)
        # --------------------------

        # 如果循环正常结束但没触发break,说明超时
        else:
            print(f"任务{task_id}超过{timeout_seconds}秒未完成,准备终止")

    except Exception as e:
        print(f"任务{task_id}出错: {str(e)}")
    finally:
        # 先尝试正常关闭浏览器
        if driver:
            try:
                driver.quit()
            except:
                pass
        # 确保Chrome进程被彻底杀掉(防止quit失败)
        if chrome_process and chrome_process.is_running():
            try:
                # 杀掉Chrome的所有子进程+主进程
                for child in chrome_process.children(recursive=True):
                    child.kill()
                chrome_process.kill()
                print(f"任务{task_id}的Chrome进程已强制终止")
            except:
                pass

# 启动并监控所有线程
def run_threads(total_threads, timeout_seconds):
    threads = []
    # 启动所有线程
    for i in range(total_threads):
        thread = threading.Thread(target=browser_task, args=(i+1, timeout_seconds))
        threads.append(thread)
        thread.start()
        print(f"启动线程{i+1}")

    # 等待线程完成,超时后确认进程状态
    for thread in threads:
        # 多给5秒处理收尾工作
        thread.join(timeout_seconds + 5)
        if thread.is_alive():
            print(f"线程{threads.index(thread)+1}仍在运行,已触发进程查杀")

if __name__ == "__main__":
    # 配置参数:修改这里的线程数和超时时间
    TOTAL_THREADS = 50
    TIMEOUT_SECONDS = 300  # 比如设置5分钟(300秒)超时
    run_threads(TOTAL_THREADS, TIMEOUT_SECONDS)

3. 关键细节说明

  • 进程绑定:通过driver.service.process.pid拿到Chrome主进程ID,用psutil杀掉整个进程树,确保浏览器彻底关闭,不会残留后台进程。
  • 超时判断:任务里通过time.time() - start_time实时计算运行时长,超过设定值就触发终止逻辑。
  • 异常处理:不管任务成功、失败还是超时,都会执行finally块里的清理操作,避免资源泄漏。

额外优化建议

  • 降低并发数:50个Chrome同时跑对CPU、内存消耗极大,建议先降到20-30个测试,看卡顿是否缓解。
  • 版本匹配:确保Chrome浏览器和ChromeDriver版本完全一致,版本不兼容是常见卡顿原因。
  • 清理缓存:在Chrome启动参数里加--disk-cache-size=0禁用缓存,减少磁盘IO占用。

内容的提问来源于stack exchange,提问作者ManuelBerate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 23:42:11