You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫爬取页面时无法自动停止的问题求助

解决Scrapy爬虫无法自动停止的问题

针对你遇到的「最后一页后网站返回重复内容、硬编码停止不灵活」的问题,以下是几个实用的解决方案:

方案1:检查分页控件的下一页按钮是否存在

多数网站的分页栏里,最后一页不会显示可点击的「下一页」按钮,你可以通过定位这个元素来判断是否继续爬取,完全不用硬编码页码。

修改代码示例:

import scrapy
from scrapy.exceptions import CloseSpider

class TowardsSustainabilitySpider(scrapy.Spider):
    name = "towards_sustainability"
    allowed_domains = ["towardssustainability.be"]
    start_urls = ["https://towardssustainability.be/products?page=1"]

    def parse(self, response):
        rows = response.xpath('//a[@class="Product-item"]')
        for row in rows:
            fund_name = row.xpath('./div/h2/text()').get()
            yield {
                'fund_name': fund_name
            }

        # 定位下一页链接(需根据网站实际分页结构调整XPath)
        # 示例:假设下一页按钮带"pagination-next"类名
        next_page = response.xpath('//a[contains(@class, "pagination-next")]/@href').get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)
        else:
            # 无下一页时停止爬虫
            raise CloseSpider(reason="No more valid pages")

提示:打开目标网站最后一页的开发者工具,查看分页区域的HTML结构,调整XPath确保能精准定位下一页按钮。

方案2:对比前后页内容的哈希值

如果网站没有明显的分页控件,可以通过计算当前页核心内容的哈希值,和上一页对比——哈希值相同说明内容重复,直接停止爬虫。

修改代码示例:

import scrapy
import hashlib
from scrapy.exceptions import CloseSpider

class TowardsSustainabilitySpider(scrapy.Spider):
    name = "towards_sustainability"
    allowed_domains = ["towardssustainability.be"]
    start_urls = ["https://towardssustainability.be/products?page=1"]
    previous_hash = None

    def parse(self, response):
        # 定位包含所有产品的父容器,生成当前页内容哈希
        product_container = response.xpath('//div[contains(@class, "ProductList")]').get()
        current_hash = hashlib.md5(product_container.encode('utf-8')).hexdigest()

        # 哈希值重复则停止
        if self.previous_hash == current_hash:
            raise CloseSpider(reason="Duplicate page content detected")
        self.previous_hash = current_hash

        # 提取数据逻辑不变
        rows = response.xpath('//a[@class="Product-item"]')
        for row in rows:
            fund_name = row.xpath('./div/h2/text()').get()
            yield {
                'fund_name': fund_name
            }

        # 继续爬取下一页
        current_page = int(response.url.split('page=')[1])
        next_page = f'https://towardssustainability.be/products?page={current_page + 1}'
        yield response.follow(next_page, callback=self.parse)

提示:调整product_container的XPath,确保只包含产品列表区域,避免页面其他动态元素(如广告、时间戳)干扰哈希值计算。

方案3:修正总页数计算逻辑

你提到第一页有结果总数但计算无效,大概率是提取总数的XPath错误或者计算逻辑有问题,重新实现这个方法:

修改代码示例:

import scrapy
import math
from scrapy.exceptions import CloseSpider

class TowardsSustainabilitySpider(scrapy.Spider):
    name = "towards_sustainability"
    allowed_domains = ["towardssustainability.be"]
    start_urls = ["https://towardssustainability.be/products?page=1"]
    total_pages = None
    items_per_page = 10

    def parse(self, response):
        # 第一页提取总条数并计算总页数
        if not self.total_pages:
            # 需根据网站实际显示调整XPath,比如找到类似"共745条"的文本
            total_text = response.xpath('//span[contains(text(), "Total")]/text()').get()
            if total_text:
                # 提取数字部分(处理千分位逗号)
                total_items = int(total_text.strip().replace(',', '').split()[-1])
                self.total_pages = math.ceil(total_items / self.items_per_page)
            else:
                # 提取失败则无限爬取(兜底用)
                self.total_pages = float('inf')

        # 提取数据逻辑不变
        rows = response.xpath('//a[@class="Product-item"]')
        for row in rows:
            fund_name = row.xpath('./div/h2/text()').get()
            yield {
                'fund_name': fund_name
            }

        # 判断是否继续爬取
        current_page = int(response.url.split('page=')[1])
        if current_page < self.total_pages:
            next_page = f'https://towardssustainability.be/products?page={current_page + 1}'
            yield response.follow(next_page, callback=self.parse)
        else:
            raise CloseSpider(reason="Reached calculated total pages")

提示:打开网站第一页,找到显示总条数的元素,复制其XPath替换示例中的对应部分,确保能正确提取数字。

内容的提问来源于stack exchange,提问作者Alexis Raphin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.30 22:09:24