Scrapy爬虫爬取页面时无法自动停止的问题求助
解决Scrapy爬虫无法自动停止的问题
针对你遇到的「最后一页后网站返回重复内容、硬编码停止不灵活」的问题,以下是几个实用的解决方案:
方案1:检查分页控件的下一页按钮是否存在
多数网站的分页栏里,最后一页不会显示可点击的「下一页」按钮,你可以通过定位这个元素来判断是否继续爬取,完全不用硬编码页码。
修改代码示例:
import scrapy from scrapy.exceptions import CloseSpider class TowardsSustainabilitySpider(scrapy.Spider): name = "towards_sustainability" allowed_domains = ["towardssustainability.be"] start_urls = ["https://towardssustainability.be/products?page=1"] def parse(self, response): rows = response.xpath('//a[@class="Product-item"]') for row in rows: fund_name = row.xpath('./div/h2/text()').get() yield { 'fund_name': fund_name } # 定位下一页链接(需根据网站实际分页结构调整XPath) # 示例:假设下一页按钮带"pagination-next"类名 next_page = response.xpath('//a[contains(@class, "pagination-next")]/@href').get() if next_page: yield response.follow(next_page, callback=self.parse) else: # 无下一页时停止爬虫 raise CloseSpider(reason="No more valid pages")
提示:打开目标网站最后一页的开发者工具,查看分页区域的HTML结构,调整XPath确保能精准定位下一页按钮。
方案2:对比前后页内容的哈希值
如果网站没有明显的分页控件,可以通过计算当前页核心内容的哈希值,和上一页对比——哈希值相同说明内容重复,直接停止爬虫。
修改代码示例:
import scrapy import hashlib from scrapy.exceptions import CloseSpider class TowardsSustainabilitySpider(scrapy.Spider): name = "towards_sustainability" allowed_domains = ["towardssustainability.be"] start_urls = ["https://towardssustainability.be/products?page=1"] previous_hash = None def parse(self, response): # 定位包含所有产品的父容器,生成当前页内容哈希 product_container = response.xpath('//div[contains(@class, "ProductList")]').get() current_hash = hashlib.md5(product_container.encode('utf-8')).hexdigest() # 哈希值重复则停止 if self.previous_hash == current_hash: raise CloseSpider(reason="Duplicate page content detected") self.previous_hash = current_hash # 提取数据逻辑不变 rows = response.xpath('//a[@class="Product-item"]') for row in rows: fund_name = row.xpath('./div/h2/text()').get() yield { 'fund_name': fund_name } # 继续爬取下一页 current_page = int(response.url.split('page=')[1]) next_page = f'https://towardssustainability.be/products?page={current_page + 1}' yield response.follow(next_page, callback=self.parse)
提示:调整product_container的XPath,确保只包含产品列表区域,避免页面其他动态元素(如广告、时间戳)干扰哈希值计算。
方案3:修正总页数计算逻辑
你提到第一页有结果总数但计算无效,大概率是提取总数的XPath错误或者计算逻辑有问题,重新实现这个方法:
修改代码示例:
import scrapy import math from scrapy.exceptions import CloseSpider class TowardsSustainabilitySpider(scrapy.Spider): name = "towards_sustainability" allowed_domains = ["towardssustainability.be"] start_urls = ["https://towardssustainability.be/products?page=1"] total_pages = None items_per_page = 10 def parse(self, response): # 第一页提取总条数并计算总页数 if not self.total_pages: # 需根据网站实际显示调整XPath,比如找到类似"共745条"的文本 total_text = response.xpath('//span[contains(text(), "Total")]/text()').get() if total_text: # 提取数字部分(处理千分位逗号) total_items = int(total_text.strip().replace(',', '').split()[-1]) self.total_pages = math.ceil(total_items / self.items_per_page) else: # 提取失败则无限爬取(兜底用) self.total_pages = float('inf') # 提取数据逻辑不变 rows = response.xpath('//a[@class="Product-item"]') for row in rows: fund_name = row.xpath('./div/h2/text()').get() yield { 'fund_name': fund_name } # 判断是否继续爬取 current_page = int(response.url.split('page=')[1]) if current_page < self.total_pages: next_page = f'https://towardssustainability.be/products?page={current_page + 1}' yield response.follow(next_page, callback=self.parse) else: raise CloseSpider(reason="Reached calculated total pages")
提示:打开网站第一页,找到显示总条数的元素,复制其XPath替换示例中的对应部分,确保能正确提取数字。
内容的提问来源于stack exchange,提问作者Alexis Raphin
相关产品推荐
相关产品推荐

