如何用Python Selenium实现Home Depot网站多页数据爬取?
实现Home Depot爬虫自动翻页抓取所有页面数据
当然可以实现自动翻页并抓取所有页面的数据,你只需要把当前的抓取逻辑放到循环中,每次抓取完当前页后尝试点击下一页按钮,直到没有下一页为止。下面是修改后的完整代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException, ElementClickInterceptedException import pandas as pd website = 'https://www.homedepot.com/b/Milwaukee/Special-Values/N-5yc1vZ7Zzv' path = '/Users/Office/Documents/chromedriver.exe' driver = webdriver.Chrome(path) driver.get(website) # 初始化存储数据的列表,放在循环外避免重复清空 prod_num = [] prod_price = [] while True: # 等待当前页SKU和价格加载完成 skus = WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'product-identifier--bd1f5'))) prices = WebDriverWait(driver, 10).until(EC.presence_of_all_elements_located((By.CLASS_NAME, 'price-format__main-price'))) # 追加当前页数据到列表 for sku in skus: prod_num.append(sku.text) for price in prices: prod_price.append(price.text) try: # 定位下一页按钮并点击 next_button = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, 'a.pagination__next')) ) # 检查按钮是否被禁用(最后一页时按钮会有disabled属性) if 'disabled' in next_button.get_attribute('class'): break next_button.click() # 等待页面跳转完成,确保新页面加载完毕 WebDriverWait(driver, 10).until( EC.staleness_of(skus[0]) # 等待旧页面元素失效,说明新页面已加载 ) except (NoSuchElementException, ElementClickInterceptedException): # 没有找到下一页按钮或者无法点击,说明到了最后一页 break driver.quit() # 保存数据到CSV df = pd.DataFrame({'code': prod_num, 'price': prod_price}) df.to_csv('HD_all_pages.csv', index=False) print(df)
关键修改说明:
- 循环逻辑:用
while True循环持续抓取,直到无法找到可点击的下一页按钮时终止 - 数据存储:把
prod_num和prod_price初始化移到循环外,确保每一页的数据都能追加到列表中,不会被清空 - 下一页判断:通过检查下一页按钮的
disabled属性,或者捕获找不到按钮的异常来判断是否到达最后一页 - 页面加载等待:使用
staleness_of等待旧页面元素失效,确保新页面完全加载后再抓取数据,避免重复抓取旧页面内容 - 异常处理:捕获
NoSuchElementException和ElementClickInterceptedException,防止因弹窗、按钮不可点击等意外情况导致程序崩溃
内容的提问来源于stack exchange,提问作者ryan houghton
相关产品推荐
相关产品推荐

