无需重复调用Selenium driver.add_cookie(),解决Scrapy爬取弹窗阻塞问题
解决Scrapy爬取时弹窗阻塞内容加载的高效方案
问题核心
爬取产品页面时,页面加载3-5秒后弹出订阅弹窗,目标HTML需关闭弹窗后才加载;直接用Selenium每个请求新建Chrome实例会导致系统负载过高,尝试传递cookie的方案因driver作用域问题失败。
高效解决方案
1. 复用单个Selenium实例
在Spider生命周期内只初始化一次Chrome实例,所有产品请求共用该实例,避免重复创建浏览器进程。同时统一处理弹窗逻辑,减少重复操作。
2. 提前设置防弹窗Cookie(可选)
如果网站通过Cookie控制弹窗展示,可以在初始化driver时设置对应Cookie,直接阻止弹窗加载,无需手动关闭。
优化后代码示例
import scrapy from selenium.webdriver import Chrome, ChromeOptions from scrapy.crawler import CrawlerProcess from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pickle class ProductSpider(scrapy.Spider): name = 'product_spider' custom_settings = { 'FEEDS': {'data/file/relevantData.jsonl': {'format': 'jsonlines', 'overwrite': True}} } def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) # 初始化单个Chrome实例 driver_path = r'C:\Users\Me\Desktop\chromedriver.exe' # 补全完整驱动路径 options = ChromeOptions() options.headless = False # 调试时可改为False,上线建议设为True # 添加参数减少不必要资源加载,提升速度 options.add_argument('--disable-images') options.add_argument('--disable-gpu') self.driver = Chrome(executable_path=driver_path, options=options) self.wait = WebDriverWait(self.driver, 10) def start_requests(self): self.driver.get('https://www.desiredurl.com/whatever') # 保存Cookie供后续复用(可选) pickle.dump(self.driver.get_cookies(), open('relevantSite_cookies.pkl', 'wb')) # 提取产品链接 try: link_elements = self.wait.until(EC.visibility_of_all_elements_located((By.XPATH, 'relevantXPATH'))) for link_el in link_elements: url = link_el.get_attribute('href') yield scrapy.Request(url, callback=self.parse) except Exception as e: self.logger.error(f"提取链接失败: {str(e)}") def parse(self, response): # 用已初始化的driver加载产品页 self.driver.get(response.url) # 处理订阅弹窗:等待弹窗出现后关闭 try: # 替换为实际弹窗关闭按钮的选择器 close_btn = self.wait.until(EC.element_to_be_clickable((By.XPATH, '//button[contains(@class, "close-modal")]'))) close_btn.click() # 等待目标产品内容加载完成 self.wait.until(EC.visibility_of_element_located((By.XPATH, '//div[@class="product-detail"]'))) except Exception as e: self.logger.warning(f"弹窗未出现或处理失败: {str(e)}") # 将driver页面转为Scrapy Response,方便用Scrapy选择器解析 page_source = self.driver.page_source scrapy_response = scrapy.http.HtmlResponse(url=response.url, body=page_source, encoding='utf-8') # 开始解析产品数据 self.logger.info('Parsing begins') # 示例:提取产品名称和价格 product_name = scrapy_response.xpath('//h1[@class="product-title"]/text()').get().strip() product_price = scrapy_response.xpath('//span[@class="price"]/text()').get().strip() yield { 'product_name': product_name, 'product_price': product_price } def closed(self, reason): # Spider结束时关闭driver,释放系统资源 self.driver.quit() self.logger.info('Chrome实例已关闭') process = CrawlerProcess() process.crawl(ProductSpider) process.start()
关键优化点
- 复用driver实例:在
__init__中初始化driver,整个Spider生命周期共用,避免重复创建浏览器进程。 - 统一弹窗处理:在
parse方法中用同一个driver处理弹窗,无需每个请求重复初始化操作。 - 资源优化:添加Chrome参数禁用图片、GPU加速,降低内存占用并提升页面加载速度。
- 优雅释放资源:通过
closed方法在Spider结束时关闭driver,避免内存泄漏。
替代方案(无需Selenium)
如果网站逻辑允许,可尝试以下无浏览器方案:
- 抓包分析弹窗加载接口,请求时携带特定参数阻止弹窗触发。
- 查找控制弹窗展示的Cookie,在Scrapy的
Request中直接携带该Cookie,跳过弹窗加载。
内容的提问来源于stack exchange,提问作者Peter Macron
相关产品推荐
相关产品推荐

