You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python类中定义可复用的Selenium驱动上下文管理器?

解决方案

直接把上下文管理器逻辑实现到你的爬虫类本身,让类成为上下文管理器,这样就能在整个程序运行期间复用同一个Selenium驱动,同时保证驱动在程序结束或异常时自动关闭,效果类似requests.Session的复用逻辑。

实现步骤与代码示例

  1. 让爬虫类实现__enter__和__exit__方法,这两个方法是上下文管理器的核心:
    • __enter__:负责初始化驱动、添加Cookie,最后返回类实例供后续调用
    • __exit__:负责在离开上下文时关闭驱动,无论是否发生异常
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from bs4 import BeautifulSoup

class SiteCrawler:
    def __init__(self):
        # 提前配置浏览器选项,比如无头模式提升速度
        self.browser_options = Options()
        self.browser_options.add_argument("--headless=new")
        self.browser_options.add_argument("--disable-gpu")
        self.browser_options.add_argument("--blink-settings=imagesEnabled=false")  # 禁用图片加载
        self.driver = None

    def __enter__(self):
        # 初始化驱动,只执行一次
        self.driver = webdriver.Chrome(options=self.browser_options)
        # 先访问目标网站域名,确保Cookie能正确添加(Cookie需要对应域名)
        self.driver.get("https://你的目标网站域名.com")
        # 批量添加Cookie示例,根据实际需求调整
        cookies = [
            {"name": "cookie1", "value": "value1", "domain": ".你的目标网站域名.com"},
            {"name": "cookie2", "value": "value2", "domain": ".你的目标网站域名.com"}
        ]
        for cookie in cookies:
            self.driver.add_cookie(cookie)
        return self  # 返回类实例,这样with块里可以直接用它调用方法

    def __exit__(self, exc_type, exc_val, exc_tb):
        # 确保驱动被关闭,即使程序崩溃
        if self.driver:
            try:
                self.driver.quit()
            except Exception:
                # 忽略关闭时的异常,避免影响程序退出
                pass

    def process_url(self, url):
        # 复用已有的驱动请求页面
        self.driver.get(url)
        # 用BeautifulSoup解析页面
        soup = BeautifulSoup(self.driver.page_source, "html.parser")
        # 这里可以添加你的页面数据提取逻辑,比如返回解析后的结果
        return soup
  1. 使用方式:
    在with块内创建爬虫实例,整个块内的所有请求都会复用同一个驱动,离开with块后驱动自动关闭。
# 在with块内使用爬虫,驱动全程保持打开
with SiteCrawler() as crawler:
    # 处理第一个页面
    page1_soup = crawler.process_url("https://你的目标网站域名.com/page1")
    # 提取page1数据...
    
    # 处理第二个页面,复用同一个驱动
    page2_soup = crawler.process_url("https://你的目标网站域名.com/page2")
    # 提取page2数据...
    
    # 甚至可以把crawler实例传给其他函数,只要不离开with块,驱动就可用
    # other_function(crawler)

# 离开with块后,驱动已经自动关闭,无需手动处理

额外优化建议

  • 可以在__enter__里添加隐式等待配置(self.driver.implicitly_wait(10)),避免元素加载超时
  • 如果需要长时间运行,可定期清理浏览器缓存,但一般爬取场景下不需要
  • 若要跨模块复用,只需在主模块的with块内把crawler实例传递给其他模块的函数即可,不要在其他模块单独创建驱动

内容的提问来源于stack exchange,提问作者TheTimebreaker

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 17:32:05