You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium+BeautifulSoup分页爬取数据重复、元素报错修复方案

多页爬取问题修复

问题现象

在Spyder环境爬取目标站点外套分类商品时,站点共6个分页(单页30条商品,总计180条),配「下一页」跳转按钮,先后实现两种爬取逻辑均未成功:

  • 第一种方案运行后得到180行数据,但所有行均为第一页内容重复,后续分页数据未被爬取
  • 第二种方案运行到下一页地址提取代码时报属性错误,仅能爬取第一页内容

原有代码问题分析

方案1

核心错误
循环内仅构造了分页url,但从未对新url发起页面请求,始终使用初始加载第一页时生成的soup对象做解析,因此每次循环提取的都是第一页的商品数据,最终结果全为第一页重复内容。另外原url中&是HTML转义字符,Python代码中直接写&即可,不需要携带amp;。
原方案1代码:

#SOLUTION 1#
from selenium import webdriver
from bs4 import BeautifulSoup
import pandas as pd
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager

driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.get('https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num=1')

# 导入页面HTML到Python环境      
soup = BeautifulSoup(driver.page_source, 'lxml')

postings = soup.find_all('div', class_ = 'isp_grid_product')

# 创建数据框
df = pd.DataFrame({'Link':[''], 'Vendor':[''],'Title':[''], 'Price':['']})

# 爬取数据
for i in range (1,7): 
    url = "https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num="+str(i)+""
    postings = soup.find_all('li', class_ = 'isp_grid_product')
    for post in postings:
        link = post.find('a', class_ = 'isp_product_image_href').get('href')
        link_full = 'https://store.unionlosangeles.com'+link
        vendor = post.find('div', class_ = 'isp_product_vendor').text.strip()
        title = post.find('div', class_ = 'isp_product_title').text.strip()
        price = post.find('div', class_ = 'isp_product_price_wrapper').text.strip()
        df = df.append({'Link':link_full, 'Vendor':vendor,'Title':title, 'Price':price}, ignore_index = True)

方案2

核心错误

  1. 下一页元素选择器错误:page-item next是li标签的class,且href属性挂载在该标签内部的<a>子标签上,直接对选中的外层div元素调用get('href')会返回None,触发属性报错
  2. 混用selenium和requests时未携带合法请求头,直接发起请求容易被站点反爬策略拦截
    原方案2代码:
### SOLUTION 2 ###

from selenium import webdriver
import requests
from bs4 import BeautifulSoup
import pandas as pd
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager


driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.get('https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num=1')

# 导入页面HTML到Python环境      
soup = BeautifulSoup(driver.page_source, 'lxml')

# 创建数据框
df = pd.DataFrame({'Link':[''], 'Vendor':[''],'Title':[''], 'Price':['']})

# 爬取数据
i = 0
while i < 6:
    
    postings = soup.find_all('li', class_ = 'isp_grid_product')

    for post in postings:
        link = post.find('a', class_ = 'isp_product_image_href').get('href')
        link_full = 'https://store.unionlosangeles.com'+link
        vendor = post.find('div', class_ = 'isp_product_vendor').text.strip()
        title = post.find('div', class_ = 'isp_product_title').text.strip()
        price = post.find('div', class_ = 'isp_product_price_wrapper').text.strip()
        df = df.append({'Link':link_full, 'Vendor':vendor,'Title':title, 'Price':price}, ignore_index = True)

    # 导入下一页HTML
    next_page = 'https://store.unionlosangeles.com'+soup.find('div', class_ = 'page-item next').get('href')
    page = requests.get(next_page)
    soup = BeautifulSoup(page.text, 'lxml')
    i += 1

修正后可运行代码

采用直接构造分页url的方式,避免下一页按钮选择器出错;使用列表暂存爬取数据,最后一次性生成DataFrame,兼容新版pandas的同时提升运行效率;每次循环重新加载对应分页、更新解析对象,保证数据不重复。

from selenium import webdriver
from bs4 import BeautifulSoup
import pandas as pd
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
import time

# 初始化浏览器
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
# 用于暂存爬取结果的列表
data_list = []

# 遍历1-6页
for page_num in range(1, 7):
    url = f"https://store.unionlosangeles.com/collections/outerwear?sort_by=creation_date&page_num={page_num}"
    driver.get(url)
    # 等待页面加载完成,可根据网络情况调整等待时间
    time.sleep(2)
    # 解析当前页源码
    soup = BeautifulSoup(driver.page_source, 'lxml')
    # 提取当前页所有商品
    postings = soup.find_all('li', class_='isp_grid_product')
    for post in postings:
        link = post.find('a', class_='isp_product_image_href').get('href')
        link_full = 'https://store.unionlosangeles.com' + link
        vendor = post.find('div', class_='isp_product_vendor').text.strip()
        title = post.find('div', class_='isp_product_title').text.strip()
        price = post.find('div', class_='isp_product_price_wrapper').text.strip()
        data_list.append({
            'Link': link_full,
            'Vendor': vendor,
            'Title': title,
            'Price': price
        })

# 关闭浏览器
driver.quit()
# 生成最终DataFrame
df = pd.DataFrame(data_list)
# 打印结果校验
print(f"总计爬取数据量:{len(df)}")
print(df.head())

如果坚持用点击下一页的逻辑,只需要把下一页地址提取的代码修正为:

next_btn = soup.find('li', class_='page-item next').find('a')
if next_btn:
    next_page = 'https://store.unionlosangeles.com' + next_btn.get('href')

即可解决属性报错问题。

内容的提问来源于stack exchange,提问作者Ivan A. Ramírez Zapata

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 21:09:18