使用Selenium点击更多商品按钮爬取bauhaus站点全量建材产品数据
解决方案
核心思路是先循环点击「更多商品」按钮直到所有商品全部加载完成,再统一提取全量商品数据,具体实现如下:
前置依赖补充
需要导入Selenium的等待相关模块,处理加载等待、元素定位问题:
from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException import time
完整实现代码
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import TimeoutException, NoSuchElementException import pandas as pd import re import time # 初始化浏览器 browser = webdriver.Chrome(r'C:\Users\KristerJens\Downloads\chromedriver_win32\chromedriver') browser.maximize_window() wait = WebDriverWait(browser, 10) browser.get('https://www.bauhaus.info/baustoffe/c/10000819') # 先处理cookie同意弹窗(不处理会遮挡按钮无法点击) try: cookie_accept_btn = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[contains(text(),"同意") or contains(text(),"Accept all")]'))) cookie_accept_btn.click() time.sleep(1) except TimeoutException: # 没有弹窗就跳过 pass # 循环点击更多商品按钮,直到按钮消失 while True: try: # 定位更多商品按钮 more_btn = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[contains(text(),"more items") or contains(text(),"更多商品")]'))) # 滚动到按钮位置,避免被遮挡 browser.execute_script("arguments[0].scrollIntoView({behavior: 'smooth', block: 'center'});", more_btn) time.sleep(0.5) # 用JS点击更稳定,避免元素被遮挡的点击报错 browser.execute_script("arguments[0].click();", more_btn) # 等待新商品加载完成:判断列表长度是否增加 old_list_len = len(browser.find_elements(By.XPATH, "//ul[@class='product-list-tiles row list-unstyled']/li")) while True: time.sleep(1) new_list_len = len(browser.find_elements(By.XPATH, "//ul[@class='product-list-tiles row list-unstyled']/li")) if new_list_len > old_list_len: break except (TimeoutException, NoSuchElementException): # 找不到更多按钮说明所有商品都加载完成了,退出循环 print("所有商品加载完毕") break # 统一提取所有商品数据 names= [] specs = [] prices = [] priceUnit = [] for li in browser.find_elements(By.XPATH, "//ul[@class='product-list-tiles row list-unstyled']/li"): try: names.append(li.find_element(By.CLASS_NAME, "product-list-tile__info__name").text) specs.append(li.find_element(By.CLASS_NAME, "product-list-tile__info__attributes").text) prices.append(li.find_element(By.CLASS_NAME, "price-tag__box").text.split('\n')[0] + "€") p = li.find_element(By.CLASS_NAME, "price-tag__sales-unit").text.split('\n')[0] priceUnit.append(p[p.find("(")+1:p.find(")")]) except Exception as e: # 个别异常商品跳过即可 continue df2 = pd.DataFrame() df2['names'] = names df2['specs'] = specs df2['prices'] = prices df2['priceUnit'] = priceUnit # 可选:导出到CSV文件 # df2.to_csv("bauhaus_建材商品数据.csv", index=False, encoding='utf-8-sig') # 关闭浏览器 browser.quit()
注意事项
- 如果按钮的XPath定位不准,可以根据你自己的元素检查结果调整按钮的定位条件
- 等待时长可以根据你的网络情况适当调整,网络慢的话可以把
WebDriverWait的10秒改成15秒 - 代码里加了异常捕获,个别加载异常的商品会直接跳过,不会导致整体爬取中断
内容的提问来源于stack exchange,提问作者awi1100
相关产品推荐
相关产品推荐

