You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy-Selenium分页爬取仅返回第一页及报错问题求助

Scrapy-Selenium分页爬取问题修复方案

问题根因分析

  • 出现的CSP报错属于站点本身Content Security Policy配置错误触发的浏览器端警告,不影响爬虫逻辑执行,无需针对性处理。
  • 仅能爬取第一页的核心原因为下一页页码提取逻辑错误:当前使用的XPath选择器//button[@class='tile-root-1uO'][1]/text()选中的是分页栏第一个页码按钮(对应数值永远为1),构造的请求永久指向第一页,无法触发后续分页爬取。

修复后的完整代码

import scrapy
from scrapy_selenium import SeleniumRequest
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC

class ComputerdealsSpider(scrapy.Spider):
    name = 'produtos'
    
    def start_requests(self):
        yield SeleniumRequest(
            url='https://naturaldaterra.com.br/hortifruti.html?page=1',
            wait_time=3,
            # 新增等待规则:商品列表加载完成后再返回响应,避免未渲染完成就提取元素
            wait_until=EC.presence_of_element_located((By.XPATH, "//div[@class='gallery-items-1IC']")),
            callback=self.parse
        )

    def parse(self, response):
        for produto in response.xpath("//div[@class='gallery-items-1IC']/div"):
            yield {
                'nome_produto': produto.xpath(".//div[@class='item-nameContainer-1kz']/span/text()").get(),
                # 单个商品仅对应一个价格的话可以把getall()换成get(),避免返回列表格式
                'valor_produto': produto.xpath(".//span[@class='itemPrice-price-1R-']/text()").get(),
            }
            
        # 修正下一页提取逻辑:先取当前激活状态的页码(即当前爬取页码)
        current_page = response.xpath("//button[contains(@class, 'tile-active')]/text()").get()
        if current_page:
            current_page_num = int(current_page.strip())
            next_page_num = current_page_num + 1
            # 校验下一页页码是否存在,避免无效请求
            next_page_exists = response.xpath(f"//button[@class='tile-root-1uO'][text()='{next_page_num}']").get()
            if next_page_exists:
                absolute_url = f"https://naturaldaterra.com.br/hortifruti.html?page={next_page_num}"
                yield SeleniumRequest(
                    url=absolute_url,
                    wait_time=3,
                    wait_until=EC.presence_of_element_located((By.XPATH, "//div[@class='gallery-items-1IC']")),
                    callback=self.parse
                )

可选优化建议

  • 可以添加最大爬取页数限制,避免站点分页规则变化导致爬虫陷入死循环
  • 提取到的商品文本可以额外调用strip()方法去除首尾空格,优化数据质量

内容的提问来源于stack exchange,提问作者Lucas Guidi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 07:15:05