Selenium+BeautifulSoup分页爬取数据重复、元素报错修复方案
多页爬取问题修复
问题现象
在Spyder环境爬取目标站点外套分类商品时,站点共6个分页(单页30条商品,总计180条),配「下一页」跳转按钮,先后实现两种爬取逻辑均未成功:
- 第一种方案运行后得到180行数据,但所有行均为第一页内容重复,后续分页数据未被爬取
- 第二种方案运行到下一页地址提取代码时报属性错误,仅能爬取第一页内容
原有代码问题分析
方案1
核心错误
循环内仅构造了分页url,但从未对新url发起页面请求,始终使用初始加载第一页时生成的soup对象做解析,因此每次循环提取的都是第一页的商品数据,最终结果全为第一页重复内容。另外原url中&是HTML转义字符,Python代码中直接写&即可,不需要携带amp;。
原方案1代码:
#SOLUTION 1# from selenium import webdriver from bs4 import BeautifulSoup import pandas as pd from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get('https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num=1') # 导入页面HTML到Python环境 soup = BeautifulSoup(driver.page_source, 'lxml') postings = soup.find_all('div', class_ = 'isp_grid_product') # 创建数据框 df = pd.DataFrame({'Link':[''], 'Vendor':[''],'Title':[''], 'Price':['']}) # 爬取数据 for i in range (1,7): url = "https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num="+str(i)+"" postings = soup.find_all('li', class_ = 'isp_grid_product') for post in postings: link = post.find('a', class_ = 'isp_product_image_href').get('href') link_full = 'https://store.unionlosangeles.com'+link vendor = post.find('div', class_ = 'isp_product_vendor').text.strip() title = post.find('div', class_ = 'isp_product_title').text.strip() price = post.find('div', class_ = 'isp_product_price_wrapper').text.strip() df = df.append({'Link':link_full, 'Vendor':vendor,'Title':title, 'Price':price}, ignore_index = True)
方案2
核心错误
- 下一页元素选择器错误:
page-item next是li标签的class,且href属性挂载在该标签内部的<a>子标签上,直接对选中的外层div元素调用get('href')会返回None,触发属性报错 - 混用selenium和requests时未携带合法请求头,直接发起请求容易被站点反爬策略拦截
原方案2代码:
### SOLUTION 2 ### from selenium import webdriver import requests from bs4 import BeautifulSoup import pandas as pd from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.get('https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num=1') # 导入页面HTML到Python环境 soup = BeautifulSoup(driver.page_source, 'lxml') # 创建数据框 df = pd.DataFrame({'Link':[''], 'Vendor':[''],'Title':[''], 'Price':['']}) # 爬取数据 i = 0 while i < 6: postings = soup.find_all('li', class_ = 'isp_grid_product') for post in postings: link = post.find('a', class_ = 'isp_product_image_href').get('href') link_full = 'https://store.unionlosangeles.com'+link vendor = post.find('div', class_ = 'isp_product_vendor').text.strip() title = post.find('div', class_ = 'isp_product_title').text.strip() price = post.find('div', class_ = 'isp_product_price_wrapper').text.strip() df = df.append({'Link':link_full, 'Vendor':vendor,'Title':title, 'Price':price}, ignore_index = True) # 导入下一页HTML next_page = 'https://store.unionlosangeles.com'+soup.find('div', class_ = 'page-item next').get('href') page = requests.get(next_page) soup = BeautifulSoup(page.text, 'lxml') i += 1
修正后可运行代码
采用直接构造分页url的方式,避免下一页按钮选择器出错;使用列表暂存爬取数据,最后一次性生成DataFrame,兼容新版pandas的同时提升运行效率;每次循环重新加载对应分页、更新解析对象,保证数据不重复。
from selenium import webdriver from bs4 import BeautifulSoup import pandas as pd from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager import time # 初始化浏览器 driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) # 用于暂存爬取结果的列表 data_list = [] # 遍历1-6页 for page_num in range(1, 7): url = f"https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num={page_num}" driver.get(url) # 等待页面加载完成,可根据网络情况调整等待时间 time.sleep(2) # 解析当前页源码 soup = BeautifulSoup(driver.page_source, 'lxml') # 提取当前页所有商品 postings = soup.find_all('li', class_='isp_grid_product') for post in postings: link = post.find('a', class_='isp_product_image_href').get('href') link_full = 'https://store.unionlosangeles.com' + link vendor = post.find('div', class_='isp_product_vendor').text.strip() title = post.find('div', class_='isp_product_title').text.strip() price = post.find('div', class_='isp_product_price_wrapper').text.strip() data_list.append({ 'Link': link_full, 'Vendor': vendor, 'Title': title, 'Price': price }) # 关闭浏览器 driver.quit() # 生成最终DataFrame df = pd.DataFrame(data_list) # 打印结果校验 print(f"总计爬取数据量:{len(df)}") print(df.head())
如果坚持用点击下一页的逻辑,只需要把下一页地址提取的代码修正为:
next_btn = soup.find('li', class_='page-item next').find('a') if next_btn: next_page = 'https://store.unionlosangeles.com' + next_btn.get('href')
即可解决属性报错问题。
内容的提问来源于stack exchange,提问作者Ivan A. Ramírez Zapata
相关产品推荐
相关产品推荐

