You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Selenium实现Home Depot网站多页数据爬取?

实现Home Depot爬虫自动翻页抓取所有页面数据

当然可以实现自动翻页并抓取所有页面的数据,你只需要把当前的抓取逻辑放到循环中,每次抓取完当前页后尝试点击下一页按钮,直到没有下一页为止。下面是修改后的完整代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException
import pandas as pd

website = 'https://www.homedepot.com/b/Milwaukee/Special-Values/N-5yc1vZ7Zzv'
path = '/Users/Office/Documents/chromedriver.exe'
driver = webdriver.Chrome(path)
driver.get(website)

# 初始化存储数据的列表,放在循环外避免重复清空
prod_num = []
prod_price = []

while True:
    # 等待当前页SKU和价格加载完成
    skus = WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'product-identifier--bd1f5')))
    prices = WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'price-format__main-price')))
    
    # 追加当前页数据到列表
    for sku in skus:
        prod_num.append(sku.text)
    for price in prices:
        prod_price.append(price.text)
    
    try:
        # 定位下一页按钮并点击
        next_button = WebDriverWait(driver, 10).until(
            EC.element_to_be_clickable((By.CSS_SELECTOR, 'a.pagination__next'))
        )
        # 检查按钮是否被禁用(最后一页时按钮会有disabled属性)
        if 'disabled' in next_button.get_attribute('class'):
            break
        next_button.click()
        # 等待页面跳转完成,确保新页面加载完毕
        WebDriverWait(driver, 10).until(
            EC.staleness_of(skus[0])  # 等待旧页面元素失效,说明新页面已加载
        )
    except (NoSuchElementException, ElementClickInterceptedException):
        # 没有找到下一页按钮或者无法点击,说明到了最后一页
        break

driver.quit()

# 保存数据到CSV
df = pd.DataFrame({'code': prod_num, 'price': prod_price})
df.to_csv('HD_all_pages.csv', index=False)
print(df)

关键修改说明:

  • 循环逻辑:用while True循环持续抓取,直到无法找到可点击的下一页按钮时终止
  • 数据存储:把prod_num和prod_price初始化移到循环外,确保每一页的数据都能追加到列表中,不会被清空
  • 下一页判断:通过检查下一页按钮的disabled属性,或者捕获找不到按钮的异常来判断是否到达最后一页
  • 页面加载等待:使用staleness_of等待旧页面元素失效,确保新页面完全加载后再抓取数据,避免重复抓取旧页面内容
  • 异常处理:捕获NoSuchElementException和ElementClickInterceptedException,防止因弹窗、按钮不可点击等意外情况导致程序崩溃

内容的提问来源于stack exchange,提问作者ryan houghton

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 01:50:29