You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium点击更多商品按钮爬取bauhaus站点全量建材产品数据

解决方案

核心思路是先循环点击「更多商品」按钮直到所有商品全部加载完成,再统一提取全量商品数据,具体实现如下:

前置依赖补充

需要导入Selenium的等待相关模块,处理加载等待、元素定位问题:

from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException
import time

完整实现代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException, NoSuchElementException
import pandas as pd
import re
import time

# 初始化浏览器
browser = webdriver.Chrome(r'C:\Users\KristerJens\Downloads\chromedriver_win32\chromedriver')
browser.maximize_window()
wait = WebDriverWait(browser, 10)

browser.get('https://www.bauhaus.info/baustoffe/c/10000819')

# 先处理cookie同意弹窗(不处理会遮挡按钮无法点击)
try:
    cookie_accept_btn = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[contains(text(),"同意") or contains(text(),"Accept all")]')))
    cookie_accept_btn.click()
    time.sleep(1)
except TimeoutException:
    # 没有弹窗就跳过
    pass

# 循环点击更多商品按钮,直到按钮消失
while True:
    try:
        # 定位更多商品按钮
        more_btn = wait.until(EC.element_to_be_clickable((By.XPATH, '//button[contains(text(),"more items") or contains(text(),"更多商品")]')))
        # 滚动到按钮位置,避免被遮挡
        browser.execute_script("arguments[0].scrollIntoView({behavior: 'smooth', block: 'center'});", more_btn)
        time.sleep(0.5)
        # 用JS点击更稳定,避免元素被遮挡的点击报错
        browser.execute_script("arguments[0].click();", more_btn)
        # 等待新商品加载完成:判断列表长度是否增加
        old_list_len = len(browser.find_elements(By.XPATH, "//ul[@class='product-list-tiles row list-unstyled']/li"))
        while True:
            time.sleep(1)
            new_list_len = len(browser.find_elements(By.XPATH, "//ul[@class='product-list-tiles row list-unstyled']/li"))
            if new_list_len > old_list_len:
                break
    except (TimeoutException, NoSuchElementException):
        # 找不到更多按钮说明所有商品都加载完成了,退出循环
        print("所有商品加载完毕")
        break

# 统一提取所有商品数据
names= []
specs = []
prices = []
priceUnit = []

for li in browser.find_elements(By.XPATH, "//ul[@class='product-list-tiles row list-unstyled']/li"):
    try:
        names.append(li.find_element(By.CLASS_NAME, "product-list-tile__info__name").text)
        specs.append(li.find_element(By.CLASS_NAME, "product-list-tile__info__attributes").text)
        prices.append(li.find_element(By.CLASS_NAME, "price-tag__box").text.split('\n')[0] + "€")
        
        p = li.find_element(By.CLASS_NAME, "price-tag__sales-unit").text.split('\n')[0]
        priceUnit.append(p[p.find("(")+1:p.find(")")])
    except Exception as e:
        # 个别异常商品跳过即可
        continue

df2 = pd.DataFrame()
df2['names'] = names
df2['specs'] = specs
df2['prices'] = prices
df2['priceUnit'] = priceUnit

# 可选:导出到CSV文件
# df2.to_csv("bauhaus_建材商品数据.csv", index=False, encoding='utf-8-sig')

# 关闭浏览器
browser.quit()

注意事项

  • 如果按钮的XPath定位不准,可以根据你自己的元素检查结果调整按钮的定位条件
  • 等待时长可以根据你的网络情况适当调整,网络慢的话可以把WebDriverWait的10秒改成15秒
  • 代码里加了异常捕获,个别加载异常的商品会直接跳过,不会导致整体爬取中断

内容的提问来源于stack exchange,提问作者awi1100

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 01:39:01