You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取:无HTML页码时的分页处理(以Mecca护肤页为例)

解决Mecca护肤页面的分页爬取问题

问题说明

需要爬取Mecca护肤页面的产品品牌、名称及成分列表并保存到CSV,但原代码仅能爬取第一页,网站每页加载36个产品,URL以offset参数递增(如?offset=36、?offset=72),分页按钮无特定类属性,无法通过常规页码标识实现分页。

实现思路

利用URL的offset参数实现分页:

  • 初始偏移量设为0,每次循环递增36
  • 持续请求对应偏移量的页面,直到页面中不再包含产品元素时终止循环
  • 同步完善CSV写入逻辑,将爬取到的结构化数据直接保存到文件

修改后的完整代码

import requests
from bs4 import BeautifulSoup
import csv

baseurl = "https://www.mecca.com.au/skin-care/"
offset = 0

# 初始化CSV文件
with open('mecca_skincare_products.csv', 'w', newline='', encoding='utf-8') as csvfile:
    fieldnames = ['品牌', '产品名称', '成分']
    writer = csv.DictWriter(csvfile, fieldnames=fieldnames)
    writer.writeheader()

    while True:
        # 构造分页请求URL
        paginated_url = f"{baseurl}?offset={offset}"
        response = requests.get(paginated_url)
        soup = BeautifulSoup(response.content, "html.parser")
        products = soup.find_all('div', class_="grid-product-info")
        
        # 当前页面无产品时,终止循环
        if not products:
            break
        
        # 收集当前页所有产品的完整链接
        product_links = []
        for item in products:
            for link in item.find_all('a', href=True):
                product_url = f"https://www.mecca.com.au{link['href']}"
                product_links.append(product_url)
        
        # 遍历产品链接,提取信息并写入CSV
        for link in product_links:
            product_res = requests.get(link)
            product_soup = BeautifulSoup(product_res.content, "html.parser")
            
            try:
                brand = product_soup.find('a', class_='css-1p371np-size5-size5-sansSerif-sansSerif-brandNameLink').text.strip()
                name = product_soup.find('span', class_='css-1noela6-paragraph-paragraph-sansSerif-sansSerif-productName').text.strip()
                # 提取成分内容
                ingredients = ""
                for div in product_soup.find_all('div'):
                    div_text = div.text.strip()
                    if div_text.startswith('Ingredients:'):
                        ingredients = div_text.split('Ingredients:', 1)[-1].strip()
                        break
                # 写入CSV
                writer.writerow({'品牌': brand, '产品名称': name, '成分': ingredients})
                print(f"已保存: {brand} - {name}")
            except AttributeError:
                # 处理页面元素缺失的异常情况
                print(f"元素缺失,跳过链接: {link}")
                continue
        
        # 偏移量递增,进入下一页
        offset += 36

关键优化点

  • 修复原代码中链接拼接重复的问题(避免生成重复路径的无效链接)
  • 新增完整的CSV写入逻辑,直接持久化爬取数据
  • 添加异常处理,防止部分页面元素缺失导致程序崩溃
  • 通过offset参数循环遍历所有分页,实现全量产品爬取

内容的提问来源于stack exchange,提问作者user22276310

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 04:43:28