You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium分页爬虫遇弹窗失效问题求助

解决Selenium爬虫因弹窗导致分页失效的问题

你的问题核心是页面弹窗(比如订阅通知、广告弹窗)干扰了分页按钮的点击,或是引发元素状态异常导致爬虫中断。以下是具体解决思路和优化后的代码:

一、自动检测并关闭弹窗

在每次处理页面内容前,先检查是否有弹窗出现,若有则立即关闭。针对Home Depot的页面,常见的是邮箱订阅弹窗,可通过定位关闭按钮实现自动处理:

  • 单独封装弹窗关闭函数,兼容无弹窗的场景(避免因无弹窗报错)
  • 使用显式等待定位弹窗元素,超时则跳过

二、优化分页逻辑与异常处理

  1. 原代码仅捕获StaleElementReferenceException,需补充NoSuchElementException、TimeoutException等常见异常,防止爬虫卡死
  2. 通过判断分页按钮的aria-disabled属性,确认是否到达最后一页,避免无效循环
  3. 替换固定time.sleep()为显式等待(如等待旧页面元素失效),确保新页面加载完成后再操作,提升稳定性

三、优化元素定位策略

部分类名可能动态变化,改用结合标签和属性的CSS选择器,降低元素定位失败的概率


优化后的完整代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import (
    StaleElementReferenceException,
    NoSuchElementException,
    TimeoutException
)
import pandas as pd

website = 'https://www.homedepot.com/b/Milwaukee/Special-Values/N-5yc1vZ7Zzv'
path = '/Users/Office/Documents/chromedriver.exe'
driver = webdriver.Chrome(path)
driver.get(website)
wait = WebDriverWait(driver, 10)  # 统一设置等待超时,可根据网络情况调整

prod_num = []
prod_price = []

def close_popup():
    """关闭页面弹窗,兼容无弹窗场景"""
    try:
        # 定位常见弹窗的关闭按钮,可根据实际弹窗调整定位
        close_btn = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[@aria-label="Close"]')))
        close_btn.click()
        print("已关闭弹窗")
    except (TimeoutException, NoSuchElementException):
        pass

while True:
    try:
        # 先处理弹窗
        close_popup()
        
        # 等待商品SKU加载完成
        skus = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'span.product-identifier--bd1f5')))
        # 等待价格加载完成
        prices = wait.until(EC.presence_of_all_elements_located((By.CSS_SELECTOR, 'span.price-format__main-price')))
        
        # 收集数据,去除多余空格
        for sku in skus:
            prod_num.append(sku.text.strip())
        for price in prices:
            prod_price.append(price.text.strip())
        
        # 处理分页:检查下一页按钮状态
        next_page = wait.until(EC.presence_of_element_located((By.XPATH, '//a[@aria-label="Next"]')))
        # 判断按钮是否禁用,到达最后一页则退出
        if next_page.get_attribute('aria-disabled') == 'true':
            print("已到最后一页,结束爬取")
            break
        
        # 点击下一页,等待页面刷新(通过旧元素失效确认)
        next_page.click()
        wait.until(EC.staleness_of(skus[0]))
        
    except StaleElementReferenceException:
        # 元素过期时重新执行循环
        continue
    except (NoSuchElementException, TimeoutException):
        # 遇到定位失败或超时,提前结束爬取
        print("爬取过程中遇到异常,提前结束")
        break

driver.quit()

# 保存数据到CSV
df = pd.DataFrame({'code': prod_num, 'price': prod_price})
df.to_csv('HD_test.csv', index=False)
print(df)

代码说明

  1. close_popup函数:专门处理弹窗,无弹窗时自动跳过,避免报错
  2. 用EC.staleness_of代替固定等待:等待旧页面元素失效,确保新页面完全加载,比time.sleep()更可靠
  3. 分页按钮状态判断:通过aria-disabled属性确认是否到达最后一页,避免无效点击
  4. 多异常捕获:覆盖常见异常场景,防止爬虫因意外情况卡死

内容的提问来源于stack exchange,提问作者ryan houghton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 02:47:03