Python Selenium爬取Lowe's:缺货按钮点击失败与404错误排查
Lowe's缺货商品爬虫修复方案
问题说明
- 用于触发缺货商品筛选的
click_show_unavailable_button函数无法正常工作 - 手动在URL后添加
?refinement=1参数会返回404错误
问题1:筛选按钮点击失效的修复
核心问题
- 原代码未先加载目标页面就调用点击函数,页面还未渲染出筛选控件
- 使用
By.CLASS_NAME定位多类名元素(如styles__StyledSVG-sc-1houmlx-0 hGNiQH icon icon-plus),By.CLASS_NAME仅支持单个类名,导致元素定位失败 - 等待条件错误:点击按钮后元素不会消失,而是状态变更,
invisibility_of_element不适用
修复措施
- 调用点击函数前先加载目标页面
- 改用
By.CSS_SELECTOR定位多类名元素(用.连接多个类名) - 使用显式等待确保元素可点击,替换错误的等待条件
问题2:URL参数404的修复
Lowe's的筛选参数是动态生成的,手动拼接refinement=1不符合网站参数规则,会触发404。正确方式是通过模拟点击界面筛选按钮来启用缺货商品显示,无需手动修改URL。
修改后的完整代码
import undetected_chromedriver as uc from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd import time def scrape_page_data(driver): # 等待商品容器加载完成 WebDriverWait(driver, 10).until(EC.presence_of_element_located((By.CLASS_NAME, 'pl'))) container = driver.find_element(By.CLASS_NAME, 'pl') # 滚动加载全部内容 for _ in range(4): driver.execute_script("window.scrollBy(0, 2000);") time.sleep(2) skus = container.find_elements(By.CLASS_NAME, 'tooltip-custom') prices = container.find_elements(By.CSS_SELECTOR, 'div.prdt-actl-pr') descriptions = container.find_elements(By.CSS_SELECTOR, '.titl-cnt.titl.brnd-desc') # 提取文本,避免空元素影响 prod_num = [sku.text.strip() for sku in skus if sku.text.strip()] prod_price = [price.text.strip() for price in prices if price.text.strip()] prod_desc = [desc.text.strip() for desc in descriptions if desc.text.strip()] return prod_num, prod_price, prod_desc def pagination(driver, url, pages=1): prod_num = [] prod_price = [] prod_desc = [] page_num = 0 for i in range(1, pages + 1): driver.get(f"{url}?offset={page_num}") current_num, current_price, current_desc = scrape_page_data(driver) prod_num.extend(current_num) prod_price.extend(current_price) prod_desc.extend(current_desc) print(f"第{i}页爬取完成:SKU数{len(current_num)},价格数{len(current_price)},描述数{len(current_desc)}") page_num += 24 time.sleep(1) return prod_num, prod_price, prod_desc def click_show_unavailable_button(driver): try: # 等待筛选面板展开按钮可点击 expand_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, 'styles__StyledSVG-sc-1houmlx-0.hGNiQH.icon.icon-plus')) ) expand_btn.click() # 等待"显示缺货商品"复选框可点击并点击 filter_checkbox = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CSS_SELECTOR, '.CheckboxWrapper-j2pgcr-0.ikAusL input')) ) if not filter_checkbox.is_selected(): filter_checkbox.click() # 等待筛选生效 time.sleep(2) print("已成功启用缺货商品显示") except Exception as e: print("启用缺货商品筛选失败:", str(e)) if __name__ == "__main__": website = 'https://www.lowes.com/pl/Drywall-joint-compound-Drywall-Building-supplies/4294858286' options = Options() # options.add_argument("--geolocation=47.8410,-122.2947") # 可选:设置地理位置 # 初始化浏览器 driver = uc.Chrome(options=options) driver.get(website) time.sleep(2) # 给页面加载留缓冲时间 # 启用缺货商品筛选 click_show_unavailable_button(driver) # 爬取指定页数数据 prod_num, prod_price, prod_desc = pagination(driver, website, pages=3) # 生成DataFrame并保存 df = pd.DataFrame({ 'code': prod_num, 'price': prod_price, 'brand': prod_desc }) df.to_csv('lowesjctest.csv', index=False) print("数据已保存到lowesjctest.csv") print(df) driver.quit()
内容的提问来源于stack exchange,提问作者ryan houghton
相关产品推荐
相关产品推荐

