You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium分页爬取循环失效,请求问题排查

Home Depot分页爬虫报错问题修复

我已经写出针对Home Depot的单页爬虫代码,能正常抓取单页商品SKU和价格,但添加分页循环遍历多页数据时,代码出现异常。

单页正常代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

website = 'https://www.homedepot.com/b/Milwaukee/Special-Values/N-5yc1vZ7Zzv'
path = '/Users/Office/Documents/chromedriver.exe'
driver = webdriver.Chrome(path)
driver.get(website)


skus = WebDriverWait(driver, 5).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'product-identifier--bd1f5')))
prices = WebDriverWait(driver, 5).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'price-format__main-price')))

prod_num = []
prod_price = []

for sku in skus:
    prod_num.append(sku.text)

for price in prices:
    prod_price.append(price.text)

driver.quit()

df = pd.DataFrame({'code': prod_num, 'price': prod_price})
df.to_csv('HD_test.csv', index=False)
print(df)

报错的分页代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd

website = 'https://www.homedepot.com/b/Milwaukee/Special-Values/N-5yc1vZ7Zzv'
path = '/Users/Office/Documents/chromedriver.exe'
driver = webdriver.Chrome(path)
driver.get(website)


skus = WebDriverWait(driver, 5).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'product-identifier--bd1f5')))
prices = WebDriverWait(driver, 5).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'price-format__main-price')))

prod_num = []
prod_price = []

for i in range(72):
    for sku in skus:
        prod_num.append(sku.text)

    for price in prices:
        prod_price.append(price.text)

    next_page = driver.find_element_by_xpath('://a[aria- label="Next"]')
    next_page.click()

driver.quit()

df = pd.DataFrame({'code': prod_num, 'price': prod_price})
df.to_csv('HD_test.csv', index=False)
print(df)

问题分析与修复

存在的问题

  1. 元素引用失效:skus和prices仅在首次加载时获取,翻页后页面DOM更新,旧元素引用已过期,会触发StaleElementReferenceException。
  2. XPath语法错误:下一页按钮的XPath写法错误,开头多了冒号,且aria-label的属性名有空格、转义符使用不当,正确路径应为//a[@aria-label="Next"]。
  3. 无翻页等待逻辑:点击下一页后直接操作,新页面元素未加载完成会报错。
  4. 固定循环次数不合理:硬编码72次循环,若实际页数不足,会因找不到下一页按钮报错。

修正后的代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, TimeoutException
import pandas as pd

website = 'https://www.homedepot.com/b/Milwaukee/Special-Values/N-5yc1vZ7Zzv'
path = '/Users/Office/Documents/chromedriver.exe'
driver = webdriver.Chrome(path)
driver.get(website)

prod_num = []
prod_price = []

while True:
    try:
        # 每次翻页后重新获取当前页的SKU和价格元素
        skus = WebDriverWait(driver, 10).until(
            EC.presence_of_all_elements_located((By.CLASS_NAME, 'product-identifier--bd1f5'))
        )
        prices = WebDriverWait(driver, 10).until(
            EC.presence_of_all_elements_located((By.CLASS_NAME, 'price-format__main-price'))
        )

        # 抓取当前页数据
        for sku in skus:
            prod_num.append(sku.text)
        for price in prices:
            prod_price.append(price.text)

        # 等待下一页按钮可点击并点击
        next_page = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.XPATH, '//a[@aria-label="Next"]'))
        )
        next_page.click()
    except (NoSuchElementException, TimeoutException):
        # 找不到下一页按钮或超时,说明已到最后一页,退出循环
        break

driver.quit()

# 保存数据
df = pd.DataFrame({'code': prod_num, 'price': prod_price})
df.to_csv('HD_test.csv', index=False)
print(df)

内容的提问来源于stack exchange,提问作者ryan houghton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 02:12:16