You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping问题:代码仅返回单页首个产品,需获取全部产品信息

问题分析与修复方案

核心问题

  1. 分类标题提取逻辑错误:category-description是页面级的分类描述,单页面通常仅存在1个,与产品列表(名称/描述/价格)zip时,因前者长度为1,导致循环仅执行一次,仅抓取第一个产品。
  2. 仅处理单个URL的分页:原代码只针对plantas-ornamentais分类做了分页遍历,其余URL完全没有产品抓取逻辑。
  3. 类名匹配问题:价格类名末尾的多余空格(price-wrapper precos-estoque-cor preco-size )可能导致BeautifulSoup无法匹配到元素。

修复后的代码

import requests
from bs4 import BeautifulSoup
from openpyxl import Workbook

# 分类URL列表
urls = [
    "/especies-de-plantas/mudas-de-plantas/plantas-ornamentais",
    "/especies-de-plantas/sementes-de-plantas",
    "/solucoes-paisagisticas/solucoes-cerca-viva",
    "/especies-de-plantas/mudas-de-plantas/paisagismo-com-bambu",
    "/especies-de-plantas/mudas-de-plantas/frutiferas",
    "/especies-de-plantas/mudas-de-plantas/especies-nativas",
    "/especies-de-plantas/livros",
    "/vasos-e-mais.html",
    "/especies-de-plantas/mudas-de-plantas/suculentas",
    "/especies-de-plantas/adubos",
    "/especies-de-plantas/mudas-de-plantas/temperos",
    "/mais-para-o-jardim"
]

wb = Workbook()
sheet = wb.active
sheet.append(["Categoria", "Nome", "Descrição", "Preço"])

# 遍历所有分类URL
for url in urls:
    base_url = "https://www.sitiodamata.com.br"
    full_url = f"{base_url}{url}"
    page = 1
    
    while True:
        # 构造分页URL(第一页无需?p参数)
        page_url = f"{full_url}?p={page}" if page > 1 else full_url
        response = requests.get(page_url)
        if response.status_code != 200:
            break
        
        soup = BeautifulSoup(response.content, 'html.parser')
        
        # 获取分类名称:优先从页面标题元素提取,失败则用URL路径作为 fallback
        category_title = soup.find(class_='page-title').text.strip() if soup.find(class_='page-title') else url.split('/')[-1].replace('.html', '')
        
        # 定位所有产品的父容器,确保每个产品的信息对应准确
        product_containers = soup.find_all(class_='product-item-info')
        
        if not product_containers:
            break  # 无产品则终止当前分类的分页遍历
        
        # 逐个提取产品信息
        for container in product_containers:
            # 产品名称
            name_elem = container.find(class_='product-item-name')
            product_name = name_elem.text.strip() if name_elem else "无名称"
            
            # 产品描述
            desc_elem = container.find(class_='descricao-item')
            product_desc = desc_elem.text.strip() if desc_elem else "无描述"
            
            # 产品价格(修复类名末尾空格问题)
            price_elem = container.find(class_='price-wrapper precos-estoque-cor preco-size')
            product_price = price_elem.text.strip() if price_elem else "无价格"
            
            # 写入Excel
            sheet.append([category_title, product_name, product_desc, product_price])
        
        # 检查下一页:存在下一页且未达最大页数则继续
        next_page = soup.find('a', class_='action next')
        if not next_page or page >= 25:
            break
        
        page += 1

# 保存完整数据
wb.save("dados_sitiodamata_completos.xlsx")

关键优化点

  • 全URL分页覆盖:对所有分类URL应用统一的分页遍历逻辑,直到无下一页或达到最大页数限制。
  • 容器化信息提取:通过产品父容器product-item-info定位每个产品,彻底避免不同信息列表长度不匹配的问题。
  • 分类标题修正:从页面标题元素提取分类名称,每个产品复用该分类,不再与产品列表错误zip。
  • 稳定性增强:增加HTTP状态码检查、空值兜底处理,修复类名匹配问题,提升代码容错性。

内容的提问来源于stack exchange,提问作者Edward Newgate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.17 01:48:04