BS4结合Selenium爬取耐克男鞋数据不全问题及纯Selenium实现方案
问题背景
最初使用BeautifulSoup(BS4)搭配Selenium开发网页爬虫,目标是爬取耐克官网男士鞋类专区的商品名称、价格信息,将结果导出为CSV文件。
代码中实现了页面自动滚动加载逻辑:通过JS获取页面滚动高度,逐次滚动到底部并等待页面加载,直到滚动高度不再变化判定为加载完成。但实际运行时页面显示共有500余款商品,程序仅能采集到约100条数据;反复核对HTML标签选择器确认配置无误后,发现Selenium返回的HTML源码会随机跳过部分商品条目。
初始实现代码如下:
from bs4 import BeautifulSoup import requests from csv import writer from selenium import webdriver import time from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager #selenium 驱动初始化 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get('https://www.nike.com/w/mens-shoes-nik1zy7ok') #页面滚动逻辑 # 获取初始滚动高度 last_height = driver.execute_script("return document.body.scrollHeight") SCROLL_PAUSE_TIME = 1 while True: # 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待页面加载 time.sleep(SCROLL_PAUSE_TIME) # 计算新的滚动高度,和上一次高度对比 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 用BS4解析Selenium获取的页面源码 soup = BeautifulSoup(driver.page_source, 'html.parser') lists = soup.find_all('div', class_='product-card__info disable-animations for--product') # 遍历HTML提取所需内容,存入shoes.csv with open('shoes.csv','w', encoding='utf8',newline='') as f: thewriter= writer(f) header=['Name','Price'] thewriter.writerow(header) for list in lists: try: name = list.find('div', class_='product-card__title').text price = list.find('div',class_='product-price css-11s12ax is--current-price').text except: print("\nList finished!") break info = [name,price] thewriter.writerow(info) print(info) # 测试其他标签匹配 soup2 = BeautifulSoup(driver.page_source, 'html.parser') lists2 = soup2.find_all('div', class_='product-card__info disable-animations for--product') # 测试写入 with open('shoes2.csv','w', encoding='utf8',newline='') as f: thewriter= writer(f) header=['Name','Price'] thewriter.writerow(header) for list in lists2: try: names = list.find('div', class_='product-card__title').text prices = list.find('div',class_='product-price is--current-price css-s56yt7').text except: print("\nList finished!") break info2 = [names,prices] thewriter.writerow(info2) print(info2)
问题解决与后续优化需求
目前已经通过纯Selenium方案解决了数据采集不全的问题,除商品名称、价格外还新增采集了商品可选配色数量字段,后续计划启用Selenium的headless无头浏览器模式降低资源占用,需要更多提升爬虫运行效率的优化建议。
纯Selenium实现代码如下:
import requests from csv import writer from selenium import webdriver import time from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By import itertools #selenium驱动初始化 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get('https://www.nike.com/w/mens-shoes-nik1zy7ok') #页面滚动逻辑 # 获取初始滚动高度 last_height = driver.execute_script("return document.body.scrollHeight") SCROLL_PAUSE_TIME = .5 while True: # 滚动到页面底部 driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") # 等待页面加载 time.sleep(SCROLL_PAUSE_TIME) # 计算新的滚动高度,和上一次高度对比 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height Wname=driver.find_elements(By.CLASS_NAME, "product-card__title") Wprice=driver.find_elements(By.CLASS_NAME, "product-card__price") Wcolor=driver.find_elements(By.CLASS_NAME, "product-card__product-count") # 分三个列表存储提取到的文本内容 name=[] price=[] color=[] for i in Wname : name.append(i.text) for i in Wprice: price.append(i.text) for i in Wcolor: color.append(i.text) # 写入CSV文件 new_list=[] with open('Menshoes.csv','w', encoding='utf8',newline='') as f: thewriter= writer(f) header=['Name','Price','#Color'] thewriter.writerow(header) for n,p,q in itertools.zip_longest(name,price,color): if n: new_list.append(n) if p: new_list.append(p) if q: new_list.append(q) info = [n,p,q] thewriter.writerow(info)
优化建议
- 替换固定休眠为显式等待:别再用
time.sleep()写死等待时长,改用Selenium自带的WebDriverWait搭配预期条件判断,等商品节点加载完成、滚动高度实际更新后再执行下一步,光这一项就能砍掉一半以上的无效等待时间,还能避免页面没加载完就操作导致的漏采问题。 - 调整滚动逻辑适配虚拟列表:耐克商品列表用了虚拟滚动机制,一次性直接滚到页面底部会导致中间区域商品还没渲染就被跳过,改成每次滚动1个视口高度,滚动后确认当前区域最后一个商品卡片渲染完成再继续下滚,从根源解决DOM节点缺失、数据采不全的问题。
- 给Chrome启动加轻量化参数:开启无头模式时用新版无头参数
--headless=new,同时追加--disable-gpu、--no-sandbox、--disable-dev-shm-usage、--disable-extensions、--blink-settings=imagesEnabled=false参数,禁用图片加载、扩展插件、不必要的系统资源调用,内存占用能降60%以上,页面加载速度提升明显。 - 减少DOM重复查询:现有代码滚动完成后分三次查询名称、价格、配色字段,会触发三次全DOM扫描,改成一次性查询所有商品卡片容器,再遍历单个容器提取内部三个字段,既降低性能消耗,也能避免三个字段列表长度不匹配需要用
zip_longest补空值的问题。 - 拦截无效网络请求:通过Chrome DevTools Protocol配置请求拦截规则,屏蔽页面里的广告、埋点统计脚本、字体、视频等和商品数据无关的请求,减少不必要的网络IO,页面加载速度能提50%以上。
- 补充之前BS4方案失效的原因:虚拟列表机制下,不在可视区域的商品节点会被临时从DOM树移除,一次性滚到底的操作会跳过大量商品的渲染环节,这些商品根本不会出现在
page_source里,和选择器写法没有关系。
内容的提问来源于stack exchange,提问作者Zero404NF
相关产品推荐
相关产品推荐

