使用Selenium+BeautifulSoup+gspread爬取Nike男装页面仅能获取65-70条数据的问题求助
Selenium+BeautifulSoup+gspread爬取Nike男装页面仅能获取65-70条数据的问题求助
我正在开发一个网页爬虫,目标是抓取Nike男装页面的全部商品数据(该页面大约有1500个商品)。我需要提取的信息包括:商品名称、是否为畅销款、服装类型、可选颜色数量以及商品价格,所有这些数据都计划通过BeautifulSoup从同一个页面中提取。
以下是我用来提取单条商品数据的代码片段:
for item in items: name = item.find('a', class_ = 'product-card__link-overlay').text.strip() try: special_tag = item.find('div', class_ = 'product-card__messaging accent--color').text.strip() except: special_tag = '/' productclass = item.find('div', class_ = 'product-card__subtitle').text.strip() colours = item.find('div', class_ = 'product-card__product-count').text.strip() try: price = item.find('div', class_ = 'product-price us__styling is--current-price css-11s12ax').text.strip() except: price = item.find('div', class_ = 'product-price is--current-price css-1ydfahe').text.strip() product = {'name':name, 'special':special_tag, 'class':productclass, 'colours':colours, 'price':price} sh.append_row([str(product['name']),str(product['special']),str(product['class']),str(product['colours']),str(product['price'])])
为了确保页面加载全部商品,我用Selenium实现了滚动加载逻辑:
time.sleep(3) previous_height = driver.execute_script('return document.body.scrollHeight') while True: driver.execute_script('window.scrollTo(0,document.body.scrollHeight);') time.sleep(3) new_height = driver.execute_script('return document.body.scrollHeight') if new_height == previous_height: page_source = driver.page_source break previous_height = new_height
但现在遇到了一个棘手的问题:提取页面源码并传入BeautifulSoup后,程序只能抓取到大约65-70个商品就停止了。我已经尝试了能想到的所有方法,也查了不少资料,但都没能解决问题。
我仔细检查过完全加载后的页面,确认所有商品的HTML类都是一致的;也给可能出错的步骤加了异常处理;另外我没有使用代理,不知道是不是Nike网站的反爬机制限制了我?
以下是我的完整代码供参考:
from bs4 import BeautifulSoup import gspread import time from selenium import webdriver from selenium.webdriver.chrome.service import Service as ChromeService from webdriver_manager.chrome import ChromeDriverManager driver = webdriver.Chrome(service=ChromeService(ChromeDriverManager().install())) gc = gspread.service_account(filename='creds.json') sh = gc.open('Nike catalog').sheet1 driver.get('https://www.nike.com/w/mens-clothing-6ymx6znik1') #Scroll program time.sleep(3) previous_height = driver.execute_script('return document.body.scrollHeight') while True: driver.execute_script('window.scrollTo(0,document.body.scrollHeight);') time.sleep(3) new_height = driver.execute_script('return document.body.scrollHeight') if new_height == previous_height: page_source = driver.page_source break previous_height = new_height #Main program baseurl='https://www.nike.com/w/mens-clothing-6ymx6znik1' headers={'User-Agent':'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36'} soup = BeautifulSoup( page_source,'lxml') items = soup.find_all('div', class_ = 'product-card__body') #HTML parser for item in items: name = item.find('a', class_ = 'product-card__link-overlay').text.strip() try: special_tag = item.find('div', class_ = 'product-card__messaging accent--color').text.strip() except: special_tag = '/' productclass = item.find('div', class_ = 'product-card__subtitle').text.strip() colours = item.find('div', class_ = 'product-card__product-count').text.strip() try: price = item.find('div', class_ = 'product-price us__styling is--current-price css-11s12ax').text.strip() except: price = item.find('div', class_ = 'product-price is--current-price css-1ydfahe').text.strip() product = {'name':name, 'special':special_tag, 'class':productclass, 'colours':colours, 'price':price} sh.append_row([str(product['name']),str(product['special']),str(product['class']),str(product['colours']),str(product['price'])])
有没有人遇到过类似的问题?或者有相关经验可以指点一下?非常感谢各位的帮助!
备注:内容来源于stack exchange,提问作者Maksim Zorić
相关产品推荐
相关产品推荐

